2023 · 28 citations · 21 references
EngineeringMachine LearningSpeech RecognitionVoxconverse Test SetData ScienceSpeaker DiarizationRobust Speech RecognitionDiarization Error RateVoice RecognitionHealth SciencesComputer ScienceDeep LearningDistant Speech RecognitionSignal ProcessingSpeech CommunicationVoiceMulti-speaker Speech RecognitionSpeech ProcessingSpeech PerceptionSpeaker Recognition
Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD) method that can efficiently address the joint modeling of large-scale speakers and predict high-resolution voice activities. Experimental results show that larger speaker capacity and higher output resolution can significantly reduce the diarization error rate (DER), which achieves the new state-of-the-art performance of 4.55% on the VoxConverse test set and 10.77% on Track 1 of the DIHARD-III evaluation set under the widely-used evaluation metrics.
21
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics) · 2023 · 73.5K citations · Full text
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey et al. · 2015 · 5.7K citations
X-Vectors: Robust DNN Embeddings for Speaker Recognition
David Snyder, Daniel Garcia-Romero, Gregory Sell et al. · 2018 · 2.6K citations
Conformer: Convolution-augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung‐Cheng Chiu et al. · 2020 · 2.5K citations