British Machine Vision Conference · 2013 · 49 citations · 26 references
American Deaf CultureMultiple Instance LearningEngineeringMachine LearningCorrespondence Search SpaceLanguage LearningSpeech RecognitionNatural Language ProcessingPattern RecognitionLanguage AcquisitionMultimodal InteractionLanguage StudiesGesture ProcessingAmerican Sign LanguageCognitive ScienceMachine VisionFeature LearningMultimodal Signal ProcessingComputer ScienceComputer VisionSign LanguageSupervisory InformationSpeech ProcessingAmerican Sign Language LinguisticsBbc Tv BroadcastsLinguistics
The goal of this work is to automatically learn a large number of signs from sign language-interpreted TV broadcasts. We achieve this by exploiting supervisory information available in the subtitles of the broadcasts. However, this information is both weak and noisy and this leads to a challenging correspondence problem when trying to identify the temporal window of the sign. We make the following contributions: (i) we show that, somewhat counter-intuitively, mouth patterns are highly informative for isolating words in a language for the Deaf, and their co-occurrence with signing can be used to significantly reduce the correspondence search space; and (ii) we develop a multiple instance learning method using an efficient discriminative search, which determines a candidate list for the sign with both high recall and precision. We demonstrate the method on videos from BBC TV broadcasts, and achieve higher accuracy and recall than previous methods, despite using much simpler features.
26
Object recognition from local scale-invariant features
David Lowe · 1999 · 16.1K citations
Carsten Rother, Vladimir Kolmogorov, Andrew Blake · ACM Transactions on Graphics · 2004 · 5.7K citations · Full text
Learning realistic human actions from movies
Ivan Laptev, Marcin Marszałek, Cordelia Schmid et al. · 2008 · 3.5K citations · Full text