Concepedia

Publication | Closed Access

Articulatory Knowledge in the Recognition of Dysarthric Speech

103

Citations

53

References

2010

Year

TLDR

Disabled speech is incompatible with modern generative and acoustic‑only ASR models. This study investigates using theoretical and empirical vocal‑tract knowledge to label segmented and unsegmented sequences of atypical speech. The authors combine vocal‑tract knowledge models with discriminative models such as neural networks, SVMs, and CRFs, then compare their performance. The combined approach yields significant accuracy gains over baseline, and transforming the vocal‑tract space of disabled speakers before retraining produces high accuracy, suggesting applicability to assistive software.

Abstract

Disabled speech is not compatible with modern generative and acoustic-only models of speech recognition (ASR). This work considers the use of theoretical and empirical knowledge of the vocal tract for atypical speech in labeling segmented and unsegmented sequences. These combined models are compared against discriminative models such as neural networks, support vector machines, and conditional random fields. Results show significant improvements in accuracy over the baseline through the use of production knowledge. Furthermore, although the statistics of vocal tract movement do not appear to be transferable between regular and disabled speakers, transforming the space of the former given knowledge of the latter before retraining gives high accuracy. This work may be applied within components of assistive software for speakers with dysarthria.

References

YearCitations

Page 1