Applied Sciences · 2021 · 30 citations · 32 references
EngineeringMachine LearningSpeech CorpusNeurolinguisticsSpoken Language ProcessingCorpus LinguisticsAphasiabank DatasetSpeech RecognitionNatural Language ProcessingData ScienceComputational LinguisticsAphasiaLanguage StudiesMachine TranslationAphasic Speech RecognitionSpeech CommunicationSpeech TechnologySpeech AnalysisAutomatic Speech RecognitionLanguage RecognitionSpeech ProcessingSpeech InputSpeech PerceptionLinguistics
Automatic speech recognition in patients with aphasia is a challenging task for which studies have been published in a few languages. Reasonably, the systems reported in the literature within this field show significantly lower performance than those focused on transcribing non-pathological clean speech. It is mainly due to the difficulty of recognizing a more unintelligible voice, as well as due to the scarcity of annotated aphasic data. This work is mainly focused on applying novel semi-supervised learning methods to the AphasiaBank dataset in order to deal with these two major issues, reporting improvements for the English language and providing the first benchmark for the Spanish language for which less than one hour of transcribed aphasic speech was used for training. In addition, the influence of reinforcing the training and decoding processes with out-of-domain acoustic and text data is described by using different strategies and configurations to fine-tune the hyperparameters and the final recognition systems. The interesting results obtained encourage extending this technological approach to other languages and scenarios where the scarcity of annotated data to train recognition models is a challenging reality.
32
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey et al. · 2015 · 5.7K citations
Connectionist temporal classification
Alex Graves, Santiago Fernández, Faustino Gomez et al. · 2006 · 5.3K citations
Engineering, Machine Learning, Spoken Language Processing +23
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Daniel Park, William Chan, Yu Zhang et al. · 2019 · 3.4K citations · Full text
KenLM: Faster and Smaller Language Model Queries
Kenneth Heafield · 2011 · 1.1K citations
L. Lichteim · Brain · 1885 · 815 citations · Full text