2020 · 48 citations · 26 references
EngineeringMachine LearningKeyword EstimationVideo SummarizationSpoken Language ProcessingCorpus LinguisticsSpeech RecognitionNatural Language ProcessingData SciencePattern RecognitionComputational LinguisticsReal-time LanguageHealth SciencesAudio RetrievalAudio MiningWord SelectionAutomated Audio CaptioningSpeech ProcessingSpeech PerceptionLinguistics
One of the problems with automated audio captioning (AAC) is the indeterminacy in word selection corresponding to the audio event/scene.Since one acoustic event/scene can be described with several words, it results in a combinatorial explosion of possible captions and difficulty in training.To solve this problem, we propose a Transformer-based audio-captioning model with keyword estimation called TRACKE.It simultaneously solves the word-selection indeterminacy problem with the main task of AAC while executing the sub-task of acoustic event detection/acoustic scene classification (i.e., keyword estimation).TRACKE estimates keywords, which comprise a word set corresponding to audio events/scenes in the input audio, and generates the caption while referring to the estimated keywords to reduce word-selection indeterminacy.Experimental results on a public AAC dataset indicate that TRACKE achieved state-ofthe-art performance and successfully estimated both the caption and its keywords.
26
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics) · 2023 · 73.5K citations · Full text
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, Quoc V. Le · arXiv (Cornell University) · 2014 · 13.3K citations · Full text
Effective Approaches to Attention-based Neural Machine Translation
Thang Luong, Hieu Pham, Christopher D. Manning · 2015 · 8.5K citations · Full text