2003 · 15 citations · 7 references
EngineeringSpeech CorpusSpoken Language ProcessingCorpus LinguisticsText MiningSpeech RecognitionNatural Language ProcessingInformation RetrievalAnchor ModelsComputational LinguisticsSpeaker DiarizationRobust Speech RecognitionConversation AnalysisLanguage StudiesAutomatic TranscriptionLinguisticsSpeech CommunicationSpeech AnalysisAutomatic Speech RecognitionMulti-speaker Speech RecognitionSpeech ProcessingIndexing MethodSpeech ArchivesSpeaker Recognition
We present unsupervised speaker indexing combined with automatic speech recognition (ASR) for speech archives such as discussions. Our proposed indexing method is based on anchor models, by which we define a feature vector based on the similarity with speakers of a large scale speech database. Several techniques are introduced to improve discriminant ability. ASR is performed using the results of this indexing. No discussion corpus is available to train acoustic and language models. So we applied the speaker adaptation technique to the baseline acoustic model based on the indexing. We also constructed a language model by merging two models that cover different linguistic features. We achieved the speaker indexing accuracy of 93% and the significant improvement of ASR for real discussion data.
7
Clustering speakers by their voices
A. Solomonoff, Alexander Mielke, Mikkel N. Schmidt et al. · 2002 · 150 citations
Unknown-multiple speaker clustering using HMM
Jitendra Ajmera, Hervé Bourlard, Itshak Lapidot et al. · 2002 · 94 citations · Full text
Speaker indexing in large audio databases using anchor models
Douglas Sturim, D.A. Reynolds, Elliot Singer et al. · 2002 · 83 citations
Music, Audio Mining, Engineering +14