IEEE Transactions on Circuits and Systems for Video Technology · 2001 · 72 citations · 40 references
EngineeringMultimedia AnalysisVideo SummarizationVideo RetrievalContent-based Video ParsingSpeech RecognitionNatural Language ProcessingImage AnalysisInformation RetrievalPattern RecognitionVideo Content AnalysisMachine VisionVisual MappingsVideo UnderstandingComputer VisionEye TrackingVisual InformationSpeech ProcessingArtsMultimedia Search
A content-based video parsing and indexing method is presented in this paper, which analyzes both information sources (auditory and visual) and accounts for their inter-relations and synergy to extract high-level semantic information. Both frame- and object-based access to the visual information is employed. The aim of the method is to extract semantically meaningful video scenes and assign semantic label(s) to them. Due to the temporal nature of video, time has to be accounted for. Thus, time-constrained video representations and indices are generated. The current approach searches for specific types of content information relevant to the presence or absence of speakers or persons. Audio-source parsing and indexing leads to the extraction of a speaker label mapping of the source over time. Video-source parsing and indexing results in the extraction of a talking-face shot mapping over time. Integration of the audio and visual mappings constrained by interaction rules leads to higher levels of video abstraction and even partial detection of its context.
40
Teuvo Kohonen · Proceedings of the IEEE · 1990 · 8.1K citations
Artificial Intelligence, Intelligent Information Processing, Vector Quantization +18
Query by image and video content: the QBIC system
Myron Flickner, Harpreet Sawhney, W. Niblack et al. · Computer · 1995 · 3.1K citations
Digital processing of speech signals
Pattern Recognition · 1980 · 2.6K citations
Digital Processing of Speech Signals
M.H. Ackroyd · Electronics and Power · 1979 · 2.4K citations