IEEE/ACM Transactions on Audio Speech and Language Processing · 2015 · 505 citations · 48 references
Structured PredictionEngineeringMachine LearningLarge Language ModelRecurrent Neural NetworkWord EmbeddingsNatural Language ProcessingLarge Language ModelsSpeech RecognitionComputational LinguisticsNeural Network VariantsLanguage StudiesLanguage ModelsMachine TranslationLarge Ai ModelSequence ModellingComputer ScienceNeural NetworksDeep LearningSpeech ProcessingLanguage ModelingLinguistics
Language models have traditionally relied on count statistics from large corpora, but neural networks now provide substantial improvements in estimating word‑sequence probabilities, though their performance and computational complexity vary with architecture and context modeling. The study aims to compare count, feedforward, recurrent, and LSTM language models on large‑vocabulary speech recognition tasks and to evaluate how advanced rescoring algorithms can improve neural‑network performance. The authors evaluate each model using perplexity and word error rate, confirm the strong correlation between these metrics across model types, and apply advanced lattice‑rescoring algorithms to assess performance gains. The results show that perplexity and word error rate are strongly correlated across all language‑model variants, regardless of architecture.
Language models have traditionally been estimated based on relative frequencies, using count statistics that can be extracted from huge amounts of text data. More recently, it has been found that neural networks are particularly powerful at estimating probability distributions over word sequences, giving substantial improvements over state-of-the-art count models. However, the performance of neural network language models strongly depends on their architectural structure. This paper compares count models to feedforward, recurrent, and long short-term memory (LSTM) neural network variants on two large-vocabulary speech recognition tasks. We evaluate the models in terms of perplexity and word error rate, experimentally validating the strong correlation of the two quantities, which we find to hold regardless of the underlying type of the language model. Furthermore, neural networks incur an increased computational complexity compared to count models, and they differently model context dependences, often exceeding the number of words that are taken into account by count based approaches. These differences require efficient search methods for neural networks, and we analyze the potential improvements that can be obtained when applying advanced algorithms to the rescoring of word lattices on large-scale setups.
48
Sepp Hochreiter, Jürgen Schmidhuber · Neural Computation · 1997 · 93.8K citations
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget et al. · 2010 · 5.4K citations
Engineering, Spoken Language Processing, Recurrent Neural Network +16
Class-based n -gram models of natural language
Peter F. Brown, P.V. deSouza, Robert L. Mercer et al. · Computational Linguistics · 1992 · 2.9K citations