Proceedings of the ... IEEE International Conference on Acoustics, Speech, and Signal Processing · 2008 · 34 citations · 14 references
EngineeringSpeech CorpusSpoken Language ProcessingLanguage LearningText MiningSpeech RecognitionNatural Language ProcessingLanguage Model AdaptationLanguage DocumentationInformation RetrievalLanguage AdaptationComputational LinguisticsPresentation Slide InformationLanguage StudiesMachine TranslationLinguisticsSpeech CommunicationAutomatic Lecture TranscriptionKeyword ExtractionLanguage RecognitionSpeech ProcessingSpeech Translation
The paper addresses language model adaptation for automatic lecture transcription by fully exploiting presentation slide information used in the lecture. As the text in the presentation slides is small in its size and fragmentary in its content, a robust adaptation scheme is addressed by focusing on the keyword and topic information. Several methods are investigated and combined; first, global topic adaptation is conducted based on PLSA (Probabilistic Latent Semantic Analysis) using keywords appearing in all slides. Web text is also retrieved to enhance the relevant text. Then, local preference of the keywords are reflected with a cache model by referring to the slide used during each utterance. Experimental evaluations on real lectures show that the proposed method combining the global and local slide information achieves a significant improvement of recognition accuracy, especially in the detection rate of content keywords.
14
Probabilistic Latent Semantic Indexing
Thomas Hofmann · ACM SIGIR Forum · 2017 · 4K citations
Engineering, Probabilistic Variant, Latent Semantic Indexing +18
Probabilistic latent semantic indexing
Thomas Hofmann · 1999 · 3.9K citations · Full text
Verbo-motor priming in the phonetic encoding of real and non-words
1999 · 207 citations