Speech Communication · 2021 · 11 citations · 18 references
Concerns have been raised regarding performance disparity in automatic speech recognition (ASR) systems as they provide unequal transcription accuracy for different user groups defined by different attributes that include gender, dialect, and race. In this paper, we propose “equal accuracy ratio”, a novel inclusiveness measure for ASR systems that can be seamlessly integrated into the standard connectionist temporal classification (CTC) training pipeline of an end-to-end neural speech recognizer to increase the recognizer’s inclusiveness. We also create a novel multi-dialect benchmark dataset to study the inclusiveness of ASR, by combining data from existing corpora in seven dialects of English (African American, General American, Latino English, British English, Indian English, Afrikaaner English, and Xhosa English). Experiments on this multi-dialect corpus show that using the equal accuracy ratio as a regularization term along with CTC loss, succeeds in lowering the accuracy gap between user groups and reduces the recognition error rate compared with a non-regularized baseline. Experiments on additional speech corpora that have different user groups also confirm our findings.
18
librosa: Audio and Music Signal Analysis in Python
Brian McFee, Colin Raffel, Dawen Liang et al. · Proceedings of the Python in Science Conferences · 2015 · 2.8K citations · Full text
Deep Speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper et al. · arXiv (Cornell University) · 2014 · 1.5K citations · Full text
Racial disparities in automated speech recognition
Allison Koenecke, Andrew Nam, Emily Lake et al. · Proceedings of the National Academy of Sciences · 2020 · 617 citations · Full text
Fairness Beyond Disparate Treatment & Disparate Impact
Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez-Rodriguez et al. · 2017 · 590 citations · Full text