arXiv (Cornell University) · 2021 · 15 citations · 47 references
Open access
Geometric LearningEngineeringMachine LearningComputer-aided Design3D Computer VisionImage AnalysisData SciencePattern RecognitionGeometric ModelingRich 2DMachine VisionNovel 3DComputer ScienceDeep Learning3D Object RecognitionComputer VisionExpansive 3DScene UnderstandingPixel-to-point Knowledge TransferScene Modeling
Most of the 3D networks are trained from scratch owning to the lack of large-scale labeled datasets. In this paper, we present a novel 3D pretraining method by leveraging 2D networks learned from rich 2D datasets. We propose the pixel-to-point knowledge transfer to effectively utilize the 2D information by mapping the pixel-level and point-level features into the same embedding space. Due to the heterogeneous nature between 2D and 3D networks, we introduce the back-projection function to align the features between 2D and 3D to make the transfer possible. Additionally, we devise an upsampling feature projection layer to increase the spatial resolution of high-level 2D feature maps, which helps learning fine-grained 3D representations. With a pretrained 2D network, the proposed pretraining process requires no additional 2D or 3D labeled data, further alleviating the expansive 3D data annotation cost. To the best of our knowledge, we are the first to exploit existing 2D trained weights to pretrain 3D deep neural networks. Our intensive experiments show that the 3D models pretrained with 2D knowledge boost the performances across various real-world 3D downstream tasks.
47
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher et al. · 2009 IEEE Conference on Computer Vision and Pattern Recognition · 2009 · 60.2K citations
Distilling the Knowledge in a Neural Network
Geoffrey E. Hinton, Oriol Vinyals · arXiv (Cornell University) · 2015 · 13.9K citations · Full text
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala et al. · 2017 · 11.1K citations
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
Raffaelli Charles, Hao Su, Kaichun Mo et al. · 2017 · 9.6K citations
Vision meets robotics: The KITTI dataset
Andreas Geiger, Philip Lenz, Christoph Stiller et al. · The International Journal of Robotics Research · 2013 · 9.4K citations