Multitalker Speech Separation With Utterance-Level Permutation Invariant Training of Deep Recurrent Neural Networks

TLDR

uPIT is a practically applicable, end‑to‑end, deep‑learning solution for speaker‑independent multitalker speech separation. The paper proposes the utterance‑level permutation invariant training (uPIT) technique. uPIT extends PIT with an utterance‑level cost function and uses RNNs that minimize utterance‑level separation error, aligning frames of the same speaker to the same output stream and eliminating the need for permutation resolution at inference. uPIT enables RNNs to separate multitalker speech without prior knowledge of duration, speaker count, identity, or gender, outperforms NMF and CASA baselines, compares favorably with deep clustering and attractor networks, generalizes to unseen speakers and languages, and a single model handles both two‑ and three‑speaker mixtures.

Abstract

In this paper, we propose the utterance-level permutation invariant training (uPIT) technique. uPIT is a practically applicable, end-to-end, deep-learning-based solution for speaker independent multitalker speech separation. Specifically, uPIT extends the recently proposed permutation invariant training (PIT) technique with an utterance-level cost function, hence eliminating the need for solving an additional permutation problem during inference, which is otherwise required by frame-level PIT. We achieve this using recurrent neural networks (RNNs) that, during training, minimize the utterance-level separation error, hence forcing separated frames belonging to the same speaker to be aligned to the same output stream. In practice, this allows RNNs, trained with uPIT, to separate multitalker mixed speech without any prior knowledge of signal duration, number of speakers, speaker identity, or gender. We evaluated uPIT on the WSJ0 and Danish two- and three-talker mixed-speech separation tasks and found that uPIT outperforms techniques based on nonnegative matrix factorization and computational auditory scene analysis, and compares favorably with deep clustering, and the deep attractor network. Furthermore, we found that models trained with uPIT generalize well to unseen speakers and languages. Finally, we found that a single model, trained with uPIT, can handle both two-speaker, and three-speaker speech mixtures.

References

Page 1

	Year	Citations
Long Short-Term Memory Sepp Hochreiter, Jürgen Schmidhuber Neural Computation	1997	93.8K
Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups Geoffrey E. Hinton, Li Deng, Dong Yu, IEEE Signal Processing Magazine EngineeringMachine LearningAcoustic ModelingSpeech RecognitionData Science	2012	10.2K
Algorithms for Non-negative Matrix Factorization Daniel D. Lee, H. Sebastian Seung	2000	5.5K
Curriculum learning Yoshua Bengio, Jérôme Louradour, Ronan Collobert, Artificial IntelligenceModel OptimizationEngineeringMachine LearningComputational Learning Theory	2009	4.8K
Some Experiments on the Recognition of Speech, with One and with Two Ears E. Colin Cherry The Journal of the Acoustical Society of America EngineeringSpeech AnalysisPhoneticsSpeech SignalsNoise	1953	4.5K
Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition George E. Dahl, Dong Yu, Li Deng, IEEE Transactions on Audio Speech and Language Processing Large-vocabulary Speech RecognitionMachine LearningEngineeringDeep Belief NetworksPhone Recognition	2011	3.1K
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs Antony W. Rix, John G. Beerends, M. P. Hollier, EngineeringSound QualitySpeech EnhancementPerceptual EvaluationCommunication	2002	3K
Performance measurement in blind audio source separation Emmanuel Vincent, Rémi Gribonval, Cédric Févotte IEEE Transactions on Audio Speech and Language Processing Source SeparationEngineeringHealth SciencesTrue Source PartAudio Signal Processing	2006	2.9K
Deep clustering: Discriminative embeddings for segmentation and separation John R. Hershey, Zhuo Chen, Jonathan Le Roux, Source SeparationSingle-channel MixturesEngineeringMachine LearningUnsupervised Machine Learning	2016	1.4K
Factorial Hidden Markov Models Zoubin Ghahramani, Michael I. Jordan Machine Learning	1997	1.2K

Page 1