International Journal of Intelligent Systems · 2021 · 28 citations · 43 references
Convolutional Neural NetworkEngineeringMachine LearningSemantic Segmentation AccuracyImage AnalysisData SciencePattern RecognitionFusion LearningSemantic SegmentationRobot LearningDeep Learning ApproachesMachine VisionFeature LearningObject DetectionData FusionComputer EngineeringComputer ScienceDeep LearningSemantic Segmentation TechniqueFeature FusionComputer VisionMultilevel Fusion
Semantic segmentation technique plays a crucial role in Internet of Things applications, such as industrial robotics and self-driving. Recently deep learning approaches have boosted semantic segmentation accuracy greatly. However, their comprehensive performance in terms of accuracy and efficiency is still far from satisfactory. We observe that (1) accuracy-oriented methods rely on numerous convolution layers and sophisticated architectures, which result in heavy computational complexity and usually take a long time for inference; (2) efficiency-oriented methods fail to capture the multiscale context information for discriminative representations during the feature fusion process, thus leading to suboptimal performance. Previous semantic segmentation approaches fail to address these two challenges simultaneously. To tackle the dilemma of precise segmentation and efficient inference, we propose a novel lightweight Multiscale Information Fusion Network (MIFNet). Specifically, the proposed MIFNet mainly consists of two core components, that is, Pyramid Refinement Connection Module (PRCM) and Lightweight Information Fusion Module (LIFM). The PRCM exploits skip learning to establish dependency between different stages. Meanwhile, the pyramid attention mechanism (PAM) in PRCM, which adjusts the weight of hybrid pyramid attention vector to refine spatial features of low-level, is developed to alleviate the semantic gap. Moreover, the LIFM is designed to detect objects at multiple scales from the global-local perspective. In LIFM, the proposed multiscale dense concatenation (MDC) adopts various dilated convolution to extract multiscale local context information. Extensive experimental results on benchmarks data sets demonstrate the significantly better performance of the proposed MIFNet compared with most existing state-of-the-art methods.
43
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten et al. · 2017 · 43.3K citations
Geometric Learning, Convolutional Neural Network, Engineering +16
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, Trevor Darrell · 2015 · 36.2K citations
Squeeze-and-Excitation Networks
Jie Hu, Li Shen, Gang Sun · 2018 · 26.8K citations
Convolutional Neural Network, Machine Vision, Machine Learning +13
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos et al. · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2017 · 21.4K citations
Semantic Image Segmentation, Convolutional Neural Network, Scene Analysis +15