2021 · 29 citations · 21 references
Dsa ModuleMachine LearningEngineeringVideo SummarizationVideo RetrievalVideo InterpretationImage AnalysisData SciencePattern RecognitionVideo TransformerMachine VisionVideo-level Representation LearningComputer ScienceVideo UnderstandingDeep LearningComputer VisionVideo RecognitionVideo HallucinationDynamic Kernel
Long-range and short-range temporal modeling are two complementary and crucial aspects of video recognition. Most of the state-of-the-arts focus on short-range spatio-temporal modeling and then average multiple snippet-level predictions to yield the final video-level prediction. Thus, their video-level prediction does not consider spatio-temporal features of how video evolves along the temporal dimension. In this paper, we introduce a novel Dynamic Segment Aggregation (DSA) module to capture relationship among snippets. To be more specific, we attempt to generate a dynamic kernel for a convolutional operation to aggregate long-range temporal information among adjacent snippets adaptively. The DSA module is an efficient plug-and-play module and can be combined with the off-the-shelf clip-based models (i.e., TSM, I3D) to perform powerful long-range modeling with minimal overhead. The final video architecture, coined as DSANet. We conduct extensive experiments on several video recognition benchmarks (i.e., Mini-Kinetics-200, Kinetics-400, Something-Something V1 and ActivityNet) to show its superiority. Our proposed DSA module is shown to benefit various video recognition models significantly. For example, equipped with DSA modules, the top-1 accuracy of I3D ResNet-50 is improved from 74.9% to 78.2% on Kinetics-400. Codes are available at https://github.com/whwu95/DSANet.
21
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher et al. · 2009 IEEE Conference on Computer Vision and Pattern Recognition · 2009 · 60.2K citations
Squeeze-and-Excitation Networks
Jie Hu, Li Shen, Gang Sun · 2018 · 26.8K citations
Convolutional Neural Network, Machine Vision, Machine Learning +13
Xiaolong Wang, Ross Girshick, Abhinav Gupta et al. · 2018 · 11K citations
Learning Spatiotemporal Features with 3D Convolutional Networks
Du Tran, Lubomir Bourdev, Rob Fergus et al. · 2015 · 9.5K citations
Convolutional Neural Network, Image Analysis, Machine Vision +14
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama et al. · 2015 · 5.2K citations