2020 · 26 citations · 45 references
Convolutional Neural NetworkEngineeringFeature DetectionMachine LearningObject DetectorComputer ArchitectureDetection TechniqueLightweight Residual-like BackboneHardware SecurityImage AnalysisPattern RecognitionComputing SystemsComputational ImagingVideo TransformerVision RecognitionMachine VisionObject DetectionComputer EngineeringComputer ScienceDeep LearningComputer VisionHardware AccelerationCpu-only DevicesObject RecognitionTraining Strategies
Previous state-of-the-art real-time object detectors have been reported on GPUs which are extremely expensive for processing massive data and in resource-restricted scenarios. Therefore, high efficiency object detectors on CPU-only devices are urgently-needed in industry. The floating-point operations (FLOPs <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> ) of networks are not strictly proportional to the running speed on CPU devices, which inspires the design of an exactly "fast" and "accurate" object detector. After investigating the concern gaps between classification networks and detection backbones, and following the design principles of efficient networks, we propose a lightweight residual-like backbone with large receptive fields and wide dimensions for low-level features, which are crucial for detection tasks. Correspondingly, we also design a light-head detection part to match the backbone capability. Furthermore, by analyzing the drawbacks of current one-stage detector training strategies, we also propose three orthogonal training strategies-IOU-guided loss, classes-aware weighting method and balanced multitask training approach. Without bells and whistles, our proposed RefineDetLite achieves 26.8 mAP on the MSCO-CO benchmark at a speed of 130 ms/pic on a single-thread CPU. The detection accuracy can be further increased to 29.6 mAP by integrating all the proposed training strategies, without apparent speed drop.
45
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell et al. · 2014 · 31.2K citations
Convolutional Neural Network, Engineering, Machine Learning +17
Feature Pyramid Networks for Object Detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick et al. · 2017 · 27.7K citations
Feature Pyramid Networks, Convolutional Neural Network, Image Analysis +14