2014 · 44 citations · 11 references
The audio component of multimedia data can be crucial for multimedia content analysis. Bag-of-audio-words (BoAW) approach is one of the most frequently used methods to represent audio content in multimedia event detection and related tasks. The method, however, has numerous criticisms, amongst which is the loss of information in the “vector quantization” step which generates word-like units. In this work, we address this issue by employing a soft quantization representation where the distance to the nearest codeword is incorporated into the model, rather than only using the nearest codeword's index as is the case with hard quantization. We explore two techniques for soft quantization and apply it to the BoAW for multimedia event detection. We find the best setup yields a 13% improvement in mean average precision, improving performance for 27 of the 30 video events.
11
Lost in quantization: Improving particular object retrieval in large scale image databases
James Philbin, Ondřej Chum, Michael Isard et al. · 2008 · 1.5K citations
Learning mid-level features for recognition
Y-Lan Boureau, Francis Bach, Yann LeCun et al. · 2010 · 1.1K citations
Jan van Gemert, Cor J. Veenman, A.W.M. Smeulders et al. · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2009 · 729 citations