arXiv (Cornell University) · 2020 · 15 citations · 40 references
Multispectral pedestrian detection is capable of adapting to insufficient illumination conditions by leveraging color-thermal modalities. On the other hand, it is still lacking of in-depth insights on how to fuse the two modalities effectively. Compared with traditional pedestrian detection, we find multispectral pedestrian detection suffers from modality imbalance problems which will hinder the optimization process of dual-modality network and depress the performance of detector. Inspired by this observation, we propose Modality Balance Network (MBNet) which facilitates the optimization process in a much more flexible and balanced manner. Firstly, we design a novel Differential Modality Aware Fusion (DMAF) module to make the two modalities complement each other. Secondly, an illumination aware feature alignment module selects complementary features according to the illumination conditions and aligns the two modality features adaptively. Extensive experimental results demonstrate MBNet outperforms the state-of-the-arts on both the challenging KAIST and CVC-14 multispectral pedestrian datasets in terms of the accuracy and the computational efficiency. Code is available at https://github.com/CalayZhou/MBNet.
40
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Feature Pyramid Networks for Object Detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick et al. · 2017 · 27.7K citations
Feature Pyramid Networks, Convolutional Neural Network, Image Analysis +14
Focal Loss for Dense Object Detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick et al. · 2017 · 24.4K citations
Image Classification, Convolutional Neural Network, Image Analysis +15
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick et al. · arXiv (Cornell University) · 2015 · 18.2K citations · Full text
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot, Yoshua Bengio · 2010 · 12.6K citations