End-to-End Waveform Utterance Enhancement for Direct Evaluation Metrics Optimization by Fully Convolutional Neural Networks

Szu‐Wei Fu, Tao-Wei Wang, Yu Tsao, Xugang Lu, Hisashi Kawai

IEEE/ACM Transactions on Audio Speech and Language Processing · 2018 · 337 citations · 65 references

DOIFull text

Open access

Concepts

TL;DR

Speech enhancement models map noisy speech to clean speech, yet training usually optimizes frame‑based MSE while evaluation relies on perception‑based metrics such as STOI, creating a mismatch that can compromise real‑world performance. The study proposes an utterance‑based fully convolutional network framework to align training objectives with evaluation metrics. The framework employs utterance‑level FCNs that exploit long‑term temporal correlations to directly optimize perception‑based objectives such as STOI. Experimental results show that the utterance‑based FCN improves STOI scores and both human and ASR intelligibility compared to MSE‑optimized baselines.

Abstract

Speech enhancement model is used to map a noisy speech to a clean speech. In the training stage, an objective function is often adopted to optimize the model parameters. However, in the existing literature, there is an inconsistency between the model optimization criterion and the evaluation criterion for the enhanced speech. For example, in measuring speech intelligibility, most of the evaluation metric is based on a short-time objective intelligibility (STOI) measure, while the frame based mean square error (MSE) between estimated and clean speech is widely used in optimizing the model. Due to the inconsistency, there is no guarantee that the trained model can provide optimal performance in applications. In this study, we propose an end-to-end utterance-based speech enhancement framework using fully convolutional neural networks (FCN) to reduce the gap between the model optimization and the evaluation criterion. Because of the utterance-based optimization, temporal correlation information of long speech segments, or even at the entire utterance level, can be considered to directly optimize perception-based objective functions. As an example, we implemented the proposed FCN enhancement framework to optimize the STOI measure. Experimental results show that the STOI of a test speech processed by the proposed approach is better than conventional MSE-optimized speech due to the consistency between the training and the evaluation targets. Moreover, by integrating the STOI into model optimization, the intelligibility of human subjects and automatic speech recognition system on the enhanced speech is also substantially improved compared to those generated based on the minimum MSE criterion.

References

65