Proceedings of the AAAI Conference on Artificial Intelligence · 2020 · 152 citations · 28 references
Video ClipsMachine LearningEngineeringAction Recognition (Movement Science)Action Recognition (Computer Vision)Novel Self-supervised MethodVideo RetrievalVideo InterpretationImage AnalysisData SciencePattern RecognitionSelf-supervised LearningMachine VisionSpatiotemporal DiagnosticsComputer ScienceVideo UnderstandingDeep LearningComputer VisionVideo AnalysisVideo Cloze ProcedureVideo Hallucination
We propose a novel self-supervised method, referred to as Video Cloze Procedure (VCP), to learn rich spatial-temporal representations. VCP first generates “blanks” by withholding video clips and then creates “options” by applying spatio-temporal operations on the withheld clips. Finally, it fills the blanks with “options” and learns representations by predicting the categories of operations applied on the clips. VCP can act as either a proxy task or a target task in self-supervised learning. As a proxy task, it converts rich self-supervised representations into video clip operations (options), which enhances the flexibility and reduces the complexity of representation learning. As a target task, it can assess learned representation models in a uniform and interpretable manner. With VCP, we train spatial-temporal representation models (3D-CNNs) and apply such models on action recognition and video retrieval tasks. Experiments on commonly used benchmarks show that the trained models outperform the state-of-the-art self-supervised models with significant margins.
28
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher et al. · 2009 IEEE Conference on Computer Vision and Pattern Recognition · 2009 · 60.2K citations
Learning Spatiotemporal Features with 3D Convolutional Networks
Du Tran, Lubomir Bourdev, Rob Fergus et al. · 2015 · 9.5K citations
Convolutional Neural Network, Image Analysis, Machine Vision +14