2014 · 57 citations · 17 references
Speech SciencesLanguage DevelopmentSpoken Language ProcessingPhonologyAcoustic ModelingDevelopmental SpeechSpeech RecognitionChild LanguageAcoustic AdaptationRobust Speech RecognitionVoice RecognitionLanguage StudiesAcoustic AnalysisHealth SciencesPronunciation ModelingSpeech PerceptionSpeech CommunicationSpeech TechnologyVoiceSpeech AcousticsSpeech ProcessingSpeech InputAcoustic VariabilityLinguistics
Developing a robust Automatic Speech Recognition (ASR) system for children is a challenging task because of increased variability in acoustic and linguistic correlates as function of young age. The acoustic variability is mainly due to the developmental changes associated with vocal tract growth. On the linguistic side, the variability is associated with limited knowledge of vocabulary, pronunciations and other linguistic constructs. This paper presents a preliminary study towards better acoustic modeling, pronunciation modeling and front-end processing for children’s speech. Results are presented as a function of age. Speaker adaptation significantly reduces mismatch and variability improving recognition results across age groups. In addition, introduction of pronunciation modeling shows promising performance improvements.
17
Kaldi Speech Recognition Toolkit
Daniel Povey · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 2024 · 4.9K citations · Full text