2020 · 58 citations · 14 references
Input Mel Spectrogram.vocganMachine LearningEngineeringGenerative Adversarial NetworkSpeech WaveformsSpeech SynthesisComputer EngineeringHigh-fidelity Real-time VocoderOutput Waveform.vocganSpeech OutputSpeech ProcessingSynthetic Image GenerationComputer ScienceGenerative AiDeep LearningSpeech Recognition
We present a novel high-fidelity real-time neural vocoder called VocGAN.A recently developed GAN-based vocoder, MelGAN, produces speech waveforms in real-time.However, it often produces a waveform that is insufficient in quality or inconsistent with acoustic characteristics of the input mel spectrogram.VocGAN is nearly as fast as MelGAN, but it significantly improves the quality and consistency of the output waveform.VocGAN applies a multi-scale waveform generator and a hierarchically-nested discriminator to learn multiple levels of acoustic properties in a balanced way.It also applies the joint conditional and unconditional objective, which has shown successful results in high-resolution image synthesis.In experiments, VocGAN synthesizes speech waveforms 416.7x faster on a GTX 1080Ti GPU and 3.24x faster on a CPU than realtime.Compared with MelGAN, it also exhibits significantly improved quality in multiple evaluation metrics including mean opinion score (MOS) with minimal additional overhead.Additionally, compared with Parallel WaveGAN, another recently developed high-fidelity vocoder, VocGAN is 6.98x faster on a CPU and exhibits higher MOS.
14
Least Squares Generative Adversarial Networks
Xudong Mao, Qing Li, Haoran Xie et al. · 2017 · 5.1K citations
High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu et al. · 2018 · 4.3K citations
WaveNet: A Generative Model for Raw Audio
Aäron van den Oord, Sander Dieleman, Heiga Zen et al. · arXiv (Cornell University) · 2016 · 3.6K citations · Full text
Music, Engineering, Machine Learning +15