2008 · 31 citations · 12 references
Emotional Speech SynthesisVoice QualitySpeech CodingVoiceOverall Voice QualityHealth SciencesSpeech AnalysisSpeech SynthesisSpeech OutputSpeech ProcessingVoice RecognitionNeutral SpeechSpeech PerceptionVoice TechnologyVoice Conversion MethodsSpeech CommunicationSpeech TechnologySpeech Recognition
This paper presents a comparison of methods for transforming voice quality in neutral synthetic speech to match cheerful, aggressive, and depressed expressive styles. Neutral speech is generated using the unit selection system in the MARY TTS platform and a large neutral database in German. The output is modified using voice conversion techniques to match the target expressive styles, the focus being on spectral envelope conversion for transforming the overall voice quality. Various improvements over the state-of-the-art weighted codebook mapping and GMM based voice conversion frameworks are employed resulting in three algorithms. Objective evaluation results show that all three methods result in comparable reduction in objective distance to target expressive TTS outputs whereas weighted frame mapping and GMM based transformations were perceived slightly better than the weighted codebook mapping outputs in generating the target expressive style in a listening test.
12
Spectral voice conversion for text-to-speech synthesis
Alexander Kain, Michael W. Macon · 2002 · 560 citations
Voice conversion through vector quantization
Masanobu Abe, Satoshi Nakamura, Kiyohiro Shikano et al. · 2003 · 416 citations