Publication | Closed Access
A comparative study of adaptive, automatic recognition of disordered speech
118
Citations
0
References
2012
Year
Unknown Venue
Speech AsrPathological SpeechDysarthric SpeakersSpeech RecognitionPattern RecognitionPhoneticsRobust Speech RecognitionAutomatic RecognitionLanguage StudiesHealth SciencesAssistive TechnologyRehabilitationComparative StudySpeech AnalysisSpeech CommunicationSpeech TechnologySpeech-driven Assistive TechnologySpeech ProcessingSpeech InputSpeech PerceptionAtypical SpeechLinguistics
Speech‑driven assistive technology offers an alternative interface for people with physical disabilities, but dysarthric speech often hampers performance of off‑the‑shelf ASR systems, and the scarcity of suitable data makes this research area challenging. This study investigates how far fundamental training and adaptation techniques from the LVCSR community can improve ASR for dysarthric speakers. Experiments use the UAspeech database, one of the largest dysarthric corpora, to train and adapt ASR systems with maximum likelihood and MAP strategies. All adapted systems yield significant improvements over baseline, with the best models achieving an average 34 % relative gain, and the optimal configuration depends more critically on speaker severity for severely dysarthric speakers.
Speech-driven assistive technology can be an attractive alternative to conventional interfaces for people with physical disabilities. However, often the lack of motor-control of the speech articulators results in disordered speech, as condition known as dysarthria. Dysarthric speakers can generally not obtain satisfactory performances with off-the-shelf automatic speech recognition (ASR) products and disordered speech ASR is an increasingly active research area. Sparseness of suitable data is a big challenge. The experiments described here use UAspeech, one of the largest dysarthric databases available, which is still easily an order of magnitude smaller than typical speech databases. This study investigates how far fundamental training and adaptation techniques developed in the LVCSR community can take us. A variety of ASR systems using maximum likelihood and MAP adaptation strategies are established with all speakers obtaining significant improvements compared to the baseline system regardless of the severity of their condition. The best systems show on average 34% relative improvement on known published results. An analysis of the correlation between intelligibility of the speaker and the type of system which would represent an optimal operating point in terms of performance shows that for severely dysarthric speakers, the exact choice of system configuration is more critical than for speakers with less disordered speech.