2016 · 27 citations · 21 references
EngineeringMachine LearningSpeech RecognitionNatural Language ProcessingData ScienceSpeaker IdentificationSpeaker DiarizationRobust Speech RecognitionVoice RecognitionJfa SystemsHealth SciencesSpeech UtteranceComputer ScienceDeep LearningDeep Neural NetworkSpeech CommunicationMulti-speaker Speech RecognitionJoint Factor AnalysisSpeech ProcessingSpeech PerceptionLinguisticsSpeaker Recognition
The i-vector and Joint Factor Analysis (JFA) systems for text-dependent speaker verification use sufficient statistics computed from a speech utterance to estimate speaker models. These statistics average the acoustic information over the utterance thereby losing all the sequence information. In this paper, we study explicit content matching using Dynamic Time Warping (DTW) and present the best achievable error rates for speaker-dependent and speaker-independent content matching. For this purpose, a Deep Neural Network/Hidden Markov Model Automatic Speech Recognition (DNN/HMM ASR) system is used to extract content-related posterior probabilities. This approach outperforms systems using Gaussian mixture model posteriors by at least 50% Equal Error Rate (EER) on the RSR2015 in content mismatch trials. DNN posteriors are also used in i-vector and JFA systems, obtaining EERs as low as 0.02%.
21
Kaldi Speech Recognition Toolkit
Daniel Povey · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 2024 · 4.9K citations · Full text
Speaker recognition: a tutorial
J.P. Campbell · Proceedings of the IEEE · 1997 · 1.6K citations · Full text