IEEE/ACM Transactions on Audio Speech and Language Processing · 2018 · 337 citations · 65 references
EngineeringMachine LearningSpeech IntelligibilitySpeech EnhancementSpeech RecognitionRobust Speech RecognitionSpeech Enhancement ModelClean SpeechHealth SciencesNoisy SpeechComputer EngineeringSpeech OutputComputer ScienceDeep LearningDistant Speech RecognitionSignal ProcessingSpeech CommunicationMulti-speaker Speech RecognitionSpeech ProcessingSpeech SeparationSpeech Perception
Speech enhancement models map noisy speech to clean speech, yet training usually optimizes frame‑based MSE while evaluation relies on perception‑based metrics such as STOI, creating a mismatch that can compromise real‑world performance. The study proposes an utterance‑based fully convolutional network framework to align training objectives with evaluation metrics. The framework employs utterance‑level FCNs that exploit long‑term temporal correlations to directly optimize perception‑based objectives such as STOI. Experimental results show that the utterance‑based FCN improves STOI scores and both human and ASR intelligibility compared to MSE‑optimized baselines.
Speech enhancement model is used to map a noisy speech to a clean speech. In the training stage, an objective function is often adopted to optimize the model parameters. However, in the existing literature, there is an inconsistency between the model optimization criterion and the evaluation criterion for the enhanced speech. For example, in measuring speech intelligibility, most of the evaluation metric is based on a short-time objective intelligibility (STOI) measure, while the frame based mean square error (MSE) between estimated and clean speech is widely used in optimizing the model. Due to the inconsistency, there is no guarantee that the trained model can provide optimal performance in applications. In this study, we propose an end-to-end utterance-based speech enhancement framework using fully convolutional neural networks (FCN) to reduce the gap between the model optimization and the evaluation criterion. Because of the utterance-based optimization, temporal correlation information of long speech segments, or even at the entire utterance level, can be considered to directly optimize perception-based objective functions. As an example, we implemented the proposed FCN enhancement framework to optimize the STOI measure. Experimental results show that the STOI of a test speech processed by the proposed approach is better than conventional MSE-optimized speech due to the consistency between the training and the evaluation targets. Moreover, by integrating the STOI into model optimization, the intelligibility of human subjects and automatic speech recognition system on the enhanced speech is also substantially improved compared to those generated based on the minimum MSE criterion.
65
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, Trevor Darrell · 2015 · 36.2K citations
Speech Recognition with Primarily Temporal Cues
Robert V. Shannon, Fan‐Gang Zeng, Vivek Kamath et al. · Science · 1995 · 3.1K citations