2018 · 42 citations · 39 references
Scene AnalysisBottom-up ApproachMachine LearningEngineeringVideo ProcessingCommunicationSemanticsVideo RetrievalImage AnalysisData ScienceAutomatic InterpretationPattern RecognitionCamera CalibrationCamera NetworkSemantic SegmentationVideo Content AnalysisLanguage StudiesMachine VisionMain Camera StreamLinguisticsComputer ScienceVideo UnderstandingDeep LearningSports GamesComputer VisionScene InterpretationVideo CommunicationEye TrackingSoccer Games
Automatic interpretation of sports games is a major challenge, especially when these sports feature complex players organizations and game phases. This paper describes a bottom-up approach based on the extraction of semantic features from the video stream of the main camera in the particular case of soccer using scene-specific techniques. In our approach, all the features, ranging from the pixel level to the game event level, have a semantic meaning. First, we design our own scene-specific deep learning semantic segmentation network and hue histogram analysis to extract pixel-level semantics for the field, players, and lines. These pixel-level semantics are then processed to compute interpretative semantic features which represent characteristics of the game in the video stream that are exploited to interpret soccer. For example, they correspond to how players are distributed in the image or the part of the field that is filmed. Finally, we show how these interpretative semantic features can be used to set up and train a semantic-based decision tree classifier for major game events with a restricted amount of training data. The main advantages of our semantic approach are that it only requires the video feed of the main camera to extract the semantic features, with no need for camera calibration, field homography, player tracking, or ball position estimation. While the automatic interpretation of sports games remains challenging, our approach allows us to achieve promising results for the semantic feature extraction and for the classification between major soccer game events such as attack, goal or goal opportunity, defense, and middle game.
39
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher et al. · 2009 IEEE Conference on Computer Vision and Pattern Recognition · 2009 · 60.2K citations
Histograms of Oriented Gradients for Human Detection
Navneet Dalal, Bill Triggs · 2005 · 31.6K citations · Full text
Kaiming He, Georgia Gkioxari, Piotr Dollár et al. · 2017 · 27.9K citations
Object Instance Segmentation, Scene Analysis, Machine Vision +13
Martin A. Fischler, Robert C. Bolles · Communications of the ACM · 1981 · 24.9K citations · Full text
Engineering, Random Sample Consensus, Sampling Technique +20