2023 · 269 citations · 42 references
Artificial IntelligenceScene AnalysisEngineeringMachine LearningMultimodal LlmImage AnalysisVisual GroundingData SciencePattern RecognitionSegmentation NetworkMachine VisionObject DetectionAdapting Segment AnythingVision Language ModelComputer ScienceDeep LearningLarge ModelsComputer VisionScene InterpretationScene UnderstandingFoundation Models
The emergence of large foundation models has advanced AI, and Segment Anything (SAM) is a prominent model for image segmentation. This study aims to apply SAM to challenging downstream tasks where it normally underperforms, such as shadow detection and camouflaged object detection. Instead of fine‑tuning SAM, the authors introduce SAM‑Adapter, which injects domain‑specific visual prompts via lightweight adapters to incorporate task‑specific knowledge. SAM‑Adapter markedly improves SAM’s performance on difficult tasks, surpassing task‑specific networks and achieving state‑of‑the‑art results on camouflaged object detection and shadow detection, with code publicly released.
The emergence of large models, also known as foundation models, has brought significant advancements to AI research. One such model is Segment Anything (SAM), which is designed for image segmentation tasks. However, as with other foundation models, our experimental findings suggest that SAM may fail or perform poorly in certain segmentation tasks, such as shadow detection and camouflaged object detection (concealed object detection). This study first paves the way for applying the large pre-trained image segmentation model SAM to these downstream tasks, even in situations where SAM performs poorly. Rather than fine-tuning the SAM network, we propose SAM-Adapter, which incorporates domain-specific information or visual prompts into the segmentation network by using simple yet effective adapters. By integrating task-specific knowledge with general knowledge learnt by the large model, SAM-Adapter can significantly elevate the performance of SAM in challenging tasks as shown in extensive experiments. We can even outperform task-specific network models and achieve state-of-the-art performance in the task we tested: camouflaged object detection, shadow detection. Our code of adapting SAM in downstream applications have been released publicly at https://github.com/tianrun-chen/SAM-Adapter-PyTorch/ and has benefited many researchers. We believe our work opens up opportunities for utilizing SAM in downstream tasks, with potential applications in various fields, including medical image processing, agriculture, remote sensing, and more.
42
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, Trevor Darrell · 2015 · 36.2K citations
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos et al. · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2017 · 21.4K citations
Semantic Image Segmentation, Convolutional Neural Network, Scene Analysis +15
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi et al. · 2017 · 15.1K citations
Convolutional Neural Network, Scene Analysis, Machine Vision +14
Alexander M. Kirillov, Eric Mintun, Nikhila Ravi et al. · 2023 · 7.8K citations