Mixed excitation for HMM-based speech synthesis

Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura

2001 · 157 citations · 10 references

Concepts

TL;DR

Previous work showed that trained HMMs can synthesize natural sounding speech, but the output typically has a vocoded quality due to the use of a traditional excitation model with a periodic impulse train or white noise. This paper aims to improve the excitation model of an HMM-based text‑to‑speech system by incorporating a mixed excitation model used in MELP to reduce synthetic quality. The mixed excitation model’s parameters are modeled by HMMs and generated during synthesis via a parameter generation algorithm. Listening tests demonstrate that the mixed excitation model significantly improves synthesized speech quality compared with the traditional excitation model.

Abstract

This paper describes improvements on the excitation model of an HMM-based text-to-speech system. In our previous work, natural sounding speech can be synthesized from trained HMMs. However, it has a typical quality of “vocoded speech” since the system uses a traditional excitation model with either a periodic impulse train or white noise. In this paper, in order to reduce the synthetic quality, a mixed excitation model used in MELP is incorporated into the system. Excitation parameters used in mixed excitation are modeled by HMMs, and generated from HMMs by a parameter generation algorithm in the synthesis phase. The result of a listening test shows that the mixed excitation model significantly improves quality of synthesized speech as compared with the traditional excitation model.

References

10