Proceedings of the ... IEEE International Conference on Acoustics, Speech, and Signal Processing · 2008 · 116 citations · 12 references
EngineeringMachine LearningBiometricsGaussian Mixture ModelsSpeech RecognitionImage AnalysisData ScienceTelephone ApplicationsPattern RecognitionSpeaker IdentificationSpeaker DiarizationRobust Speech RecognitionSupport Vector MachinesVoice RecognitionAutomatic AgeHealth SciencesComputer SciencePolynomial KernelGmm SupervectorsComputer VisionMulti-speaker Speech RecognitionSpeech ProcessingSpeech PerceptionSpeaker Recognition
This paper compares two approaches of automatic age and gender classification with 7 classes. The first approach are Gaussian mixture models (GMMs) with universal background models (UBMs), which is well known for the task of speaker identification/verification. The training is performed by the EM algorithm or MAP adaptation respectively. For the second approach for each speaker of the test and training set a GMM model is trained. The means of each model are extracted and concatenated, which results in a GMM supervector for each speaker. These supervectors are then used in a support vector machine (SVM). Three different kernels were employed for the SVM approach: a polynomial kernel (with different polynomials), an RBF kernel and a linear GMM distance kernel, based on the KL divergence. With the SVM approach we improved the recognition rate to 74% (p < 0.001) and are in the same range as humans.
12
Emotion recognition in human-computer interaction
Roddy Cowie, Ellen Douglas‐Cowie, Nicolas Tsapatsoulis et al. · IEEE Signal Processing Magazine · 2001 · 2.5K citations · Full text