2009 · 56 citations · 6 references
EngineeringSpeech CorpusCommunicationPhonologyCorpus LinguisticsSpeech RecognitionNatural Language ProcessingLanguage DocumentationComputational LinguisticsSouth AfricaLanguage EngineeringSpeech InterfaceVoice RecognitionLanguage StudiesMachine TranslationAsr Corpus DesignIndex TermsLanguage TechnologySpeech CommunicationLanguage RecognitionLanguage CorpusSpeech ProcessingSpeech InputSpeech PerceptionLinguisticsLwazi Corpus
We investigate the number of speakers and the amount of data that is required for the development of useable speakerindependent speech-recognition systems in resource-scarce languages. Our experiments employ the Lwazi corpus, which contains speech in the eleven official languages of South Africa. We find that a surprisingly small number of speakers (fewer than 50) and around 10 to 20 hours of speech per language are sufficient for the purposes of acceptable phone-based recognition. Index Terms: speech recognition, corpus design
6
Data selection for speech recognition
Yi Wu, Rong Zhang, Alexander I. Rudnicky · 2007 · 61 citations · Full text