Extensions of recurrent neural network language model

Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jaň Černocký, Sanjeev Khudanpur

2011 · 1.6K citations · 21 references

Concepts

TL;DR

The recurrent neural network language model (RNN LM) outperforms many techniques but suffers from high computational complexity. This study proposes several modifications to the RNN LM aimed at reducing computational complexity and speeding up training and testing. The authors explore parameter‑reduction strategies to make the model smaller and faster. The modified RNN achieves more than 15‑fold speedup, benefits from backpropagation through time, outperforms feedforward networks, and is smaller, faster, and more accurate than the baseline.

Abstract

We present several modifications of the original recurrent neural network language model (RNN LM).While this model has been shown to significantly outperform many competitive language modeling techniques in terms of accuracy, the remaining problem is the computational complexity. In this work, we show approaches that lead to more than 15 times speedup for both training and testing phases. Next, we show importance of using a backpropagation through time algorithm. An empirical comparison with feedforward networks is also provided. In the end, we discuss possibilities how to reduce the amount of parameters in the model. The resulting RNN model can thus be smaller, faster both during training and testing, and more accurate than the basic one.

References

21