Parameter Definition for Multilayer Perceptron Intended for Speaker Identification

Ihor Tereikovskyi, Ihor Subach, Oleh Tereikovskyi, Liudmyla Tereikovska, Volodymyr Nakonechnyi

2019 IEEE International Conference on Advanced Trends in Information Theory (ATIT) · 2019 · 11 citations · 17 references

Concepts

Abstract

The article deals with the improvement of speaker identification tools. The prospects of neural network identification tools are established. Authors show that ways of improving such identification tools are associated with the adaptation of the neural network model used to the significant conditions of the task. The authors propose to increase the efficiency of neural network identification tools by adapting the parameters of a deep neural network with direct signal propagation to such conditions of the speaker identification task as the parameters of the voice signal, the number of identified speakers and training samples. The adaptation approach providing for the experimental definition of multilayer perceptron structural parameters is developed. The identification issue under study involved: the number of speakers - 10, voice signal duration - 7. 008s, voice signal sampling frequency 16,000 Hz, quasi-stationary fragment duration-0.016 s, and the number of Mel-frequency cepstral coefficients (MFCC) of the single quasi-stationary fragment - 20. The training sample includes 100 recordings of the voice signal for 10 speakers when they read texts in English. Recording took place in studio. Each fragment of the voice signal sample is unique. The neural network model adapted to these conditions is a three-layer perceptron with 256 neurons in the first hidden layer and 80 neurons in the second hidden layer. The ReLU activation function is used for the hidden layer neurons, and the Softmax activation function is used for the output neurons. Given an acceptable level of resource intensity, the developed neural network model allows achieving identification accuracy of about 0.95, which is comparable with the most modern means of a similar purpose. The necessity of further research on developing a method of adaptation of multilayer perceptron architectural parameters to the task of integral identification of speaker emotions and personality is substantiated.

References

17