2023 · 43 citations · 42 references
Image ClassificationConvolutional Neural NetworkMachine VisionImage AnalysisFeature DetectionEngineeringPattern RecognitionObject DetectionVision RecognitionAutomatic Target RecognitionFrequency Representation IntegrationComputer ScienceDeep LearningVideo TransformerSignal ProcessingSegment ObjectsComputer VisionDifferent Frequency Components
Recent camouflaged object detection (COD) approaches have been proposed to accurately segment objects blended into surroundings. The most challenging and critical issue in COD is to find out the lines of demarcation between objects and background in the camouflage environment. Because of the similarity between the target object and the background, these lines are difficult to be found accurately. However, these are easy to be observed in different frequency components of the image. To this end, in this paper we rethink COD from the perspective of frequency components and propose a Frequency Representation Integration Network to mine informative cues from them. Specifically, we obtain high-frequency components from the original image by Laplacian pyramid-like decomposition, and then respectively send the image to a transformer-based encoder and frequency components to a tailored CNN-based Residual Frequency Array Encoder. Besides, we utilize the multi-head self-attention in transformer encoder to capture low-frequency signals, which can effectively parse the overall contextual information of camouflage scenes. We also design a Frequency Representation Reasoning Module, which progressively eliminates discrepancies between differentiated frequency representations and integrates them by modeling their point-wise relations. Moreover, to further bridge different frequency representations, we introduce the image reconstruction task to implicitly guide their integration. Sufficient experiments on three widely-used COD benchmark datasets demonstrate that our method surpasses existing state-of-the-art methods by a large margin.
42
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher et al. · 2009 IEEE Conference on Computer Vision and Pattern Recognition · 2009 · 60.2K citations
SegFormer: Simple and Efficient Design for Semantic Segmentation with\n Transformers
Enze Xie, Wenhai Wang, Zhiding Yu et al. · arXiv (Cornell University) · 2021 · 3.2K citations · Full text
PVT v2: Improved baselines with pyramid vision transformer
Wenhai Wang, Enze Xie, Xiang Li et al. · Computational Visual Media · 2022 · 2K citations · Full text
Structure-Measure: A New Way to Evaluate Foreground Maps
Deng-Ping Fan, Ming‐Ming Cheng, Yun Liu et al. · 2017 · 1.8K citations