2020 · 31 citations · 20 references
EngineeringMachine LearningTarget VoiceSpeech RecognitionNatural Language ProcessingNoisy SamplesData ScienceRobust Speech RecognitionVoice RecognitionHealth SciencesSpeech SynthesisData Efficient VoiceSpeech OutputComputer ScienceDeep LearningSpeech CommunicationTarget SpeakerVoiceDomain Adversarial TrainingMulti-speaker Speech RecognitionDomain AdaptationSpeech ProcessingSpeech Perception
Data efficient voice cloning aims at synthesizing target speaker's voice with only a few enrollment samples at hand.To this end, speaker adaptation and speaker encoding are two typical methods based on base model trained from multiple speakers.The former uses a small set of target speaker data to transfer the multi-speaker model to target speaker's voice through direct model update, while in the latter, only a few seconds of target speaker's audio directly goes through an extra speaker encoding model along with the multi-speaker model to synthesize target speaker's voice without model update.Nevertheless, the two methods need clean target speaker data.However, the samples provided by user may inevitably contain acoustic noise in real applications.It's still challenging to generating target voice with noisy data.In this paper, we study the data efficient voice cloning problem from noisy samples under the sequenceto-sequence based TTS paradigm.Specifically, we introduce domain adversarial training (DAT) to speaker adaptation and speaker encoding, which aims to disentangle noise from speechnoise mixture.Experiments show that for both speaker adaptation and encoding, the proposed approaches can consistently synthesize clean speech from noisy speaker samples, apparently outperforming the method adopting state-of-the-art speech enhancement module.
20
Laurens van der Maaten, Geoffrey E. Hinton · Journal of Machine Learning Research · 2008 · 35.7K citations
X-Vectors: Robust DNN Embeddings for Speaker Recognition
David Snyder, Daniel Garcia-Romero, Gregory Sell et al. · 2018 · 2.6K citations
Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions
Jonathan Shen, Ruoming Pang, Ron J. Weiss et al. · 2018 · 2.6K citations
Engineering, Machine Learning, Spoken Language Processing +18
Tacotron: Towards End-to-End Speech Synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton et al. · 2017 · 1.7K citations
Engineering, Machine Learning, Spoken Language Processing +20
Tutorial on Variational Autoencoders
Carl Doersch · arXiv (Cornell University) · 2016 · 1.4K citations · Full text