2023 · 19 citations · 17 references
Anomalous Sound Detection (ASD) aims to identify whether the sound emitted from a machine is anomalous or not. Most advanced methods use 2-D CNNs to extract features of normal sounds from log-mel spectrograms for ASD. However, these methods can not fully exploit temporal information of log-mel spectrograms, resulting in poor performance on some machine types. In this paper, we propose a new framework for ASD named Spectrogram-Wavegram WaveNet (SW-WaveNet), which segments the 2-D log-mel spectrogram into 1-D waveform signals of different frequency bands and combines the representation vector extracted by WaveNet from segmented log-mel spectrograms and Wavegrams, respectively. The proposed framework utilizes WaveNet’s powerful capability of modeling waveform signals to effectively extract temporal information from log-mel spectrograms and Wavegrams. Experiments on the DCASE 2020 Challenge Task 2 dataset show that our framework achieves higher average AUC scores (93.25%) and pAUC scores (87.41%) than the previous works.
17
Laurens van der Maaten, Geoffrey E. Hinton · Journal of Machine Learning Research · 2008 · 35.7K citations
WaveNet: A Generative Model for Raw Audio
Aäron van den Oord, Sander Dieleman, Heiga Zen et al. · arXiv (Cornell University) · 2016 · 3.6K citations · Full text
Music, Engineering, Machine Learning +15
Conditional Image Generation with PixelCNN Decoders
Aäron van den Oord, Nal Kalchbrenner, Oriol Vinyals et al. · arXiv (Cornell University) · 2016 · 1.6K citations · Full text
Convolutional Neural Network, Conditional Pixelcnn, Machine Vision +14