2015 · 237 citations · 28 references
EngineeringMachine LearningEgocentric VideoImage AnalysisPattern RecognitionSelf-supervised LearningRobot LearningVideo TransformerVision RecognitionImage Representations TiedCognitive ScienceMachine VisionFeature LearningVideo UnderstandingProprioceptive Motor SignalsDeep LearningComputer VisionScene UnderstandingVisual Recognition
Understanding how images of objects and scenes behave in response to specific ego-motions is a crucial aspect of proper visual development, yet existing visual learning methods are conspicuously disconnected from the physical source of their images. We propose to exploit proprioceptive motor signals to provide unsupervised regularization in convolutional neural networks to learn visual representations from egocentric video. Specifically, we enforce that our learned features exhibit equivariance, i.e, they respond predictably to transformations associated with distinct ego-motions. With three datasets, we show that our unsupervised feature learning approach significantly outperforms previous approaches on visual recognition and next-best-view prediction tasks. In the most challenging test, we show that features learned from video captured on an autonomous driving platform improve large-scale scene recognition in static images from a disjoint domain.
28
Yangqing Jia, Evan Shelhamer, Jeff Donahue et al. · 2014 · 11.1K citations
Convolutional Neural Network, Machine Vision, Machine Learning +14
Vision meets robotics: The KITTI dataset
Andreas Geiger, Philip Lenz, Christoph Stiller et al. · The International Journal of Robotics Research · 2013 · 9.4K citations
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio et al. · 2008 · 7.2K citations
Dimensionality Reduction by Learning an Invariant Mapping
Raia Hadsell, Sumit Chopra, Yann LeCun · 2006 · 5.1K citations · Full text