arXiv (Cornell University) · 2020 · 75 citations · 45 references
We present ESPnet-SE, which is designed for the quick development of speech\nenhancement and speech separation systems in a single framework, along with the\noptional downstream speech recognition module. ESPnet-SE is a new project which\nintegrates rich automatic speech recognition related models, resources and\nsystems to support and validate the proposed front-end implementation (i.e.\nspeech enhancement and separation).It is capable of processing both\nsingle-channel and multi-channel data, with various functionalities including\ndereverberation, denoising and source separation. We provide all-in-one recipes\nincluding data pre-processing, feature extraction, training and evaluation\npipelines for a wide range of benchmark datasets. This paper describes the\ndesign of the toolkit, several important functionalities, especially the speech\nrecognition integration, which differentiates ESPnet-SE from other open source\ntoolkits, and experimental results with major benchmark datasets.\n
45
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa et al. · arXiv (Cornell University) · 2019 · 16.2K citations · Full text
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey et al. · 2015 · 5.7K citations