2007 · 38 citations · 7 references
EngineeringSpeech CorpusSynthetic VoicesCorpus LinguisticsSpeech RecognitionNatural Language ProcessingData ScienceComputational LinguisticsPhoneticsVoice RecognitionLanguage StudiesAutomatic BuildingMachine TranslationSpeech SynthesisSpeech OutputSpeech CommunicationLarge Speech FileNatural Sounding VoicesSpeech ProcessingSpeech InputSpeech PerceptionLinguistics
Large multi paragraph speech databases encapsulate prosodic and contextual information beyond the sentence level which could be exploited to build natural sounding voices. This paper discusses our efforts on automatic building of synthetic voices from large multi-paragraph speech databases. We show that the primary issue of segmentation of large speech file could be addressed with modifications to forced-alignment technique and that the proposed technique is independent of the duration of the audio file. We also discuss how this framework could be extended to build a large number of voices from public domain large multi-paragraph recordings.
7
Multi-paragraph segmentation of expository text
Marti A. Hearst · 1994 · 558 citations · Full text
Natural Language Processing, Applied Linguistics, Engineering +15
Multi-Paragraph Segmentation of Expository Text
Marti A. Hearst · ArXiv.org · 1994 · 392 citations · Full text
Natural Language Processing, Applied Linguistics, Engineering +15
Constructing stylistic synthesis databases from audio books
Yong Zhao, Peng Di, Lijuan Wang et al. · 2006 · 22 citations