Exploring Strategies for Training Deep Neural Networks

TLDR

Deep neural networks can represent highly non‑linear functions, yet training them from random initialization often fails, prompting recent work on greedy layer‑wise unsupervised pre‑training with RBMs or autoencoders to initialize deep belief networks and improve optimization. This paper empirically investigates these greedy layer‑wise training algorithms to understand why they succeed. We perform experiments varying network depth, RBM input distributions, gradient‑combination strategies, and online training to assess how these factors influence performance. The experiments confirm that greedy layer‑wise pre‑training initializes weights near good local minima, acts as regularization, improves generalization, and promotes high‑level distributed representations. Hinton et al.

Abstract

Deep multi-layer neural networks have many levels of non-linearities allowing them to compactly represent highly non-linear and highly-varying functions. However, until recently it was not clear how to train such deep networks, since gradient-based optimization starting from random initialization often appears to get stuck in poor solutions. Hinton et al. recently proposed a greedy layer-wise unsupervised learning procedure relying on the training algorithm of restricted Boltzmann machines (RBM) to initialize the parameters of a deep belief network (DBN), a generative model with many layers of hidden causal variables. This was followed by the proposal of another greedy layer-wise procedure, relying on the usage of autoassociator networks. In the context of the above optimization problem, we study these algorithms empirically to better understand their success. Our experiments confirm the hypothesis that the greedy layer-wise unsupervised training strategy helps the optimization by initializing weights in a region near a good local minimum, but also implicitly acts as a sort of regularization that brings better generalization and encourages internal distributed representations that are high-level abstractions of the input. We also present a series of experiments aimed at evaluating the link between the performance of deep neural networks and practical aspects of their topology, for example, demonstrating cases where the addition of more depth helps. Finally, we empirically explore simple variants of these training algorithms, such as the use of different RBM input unit distributions, a simple way of combining gradient estimators to improve performance, as well as on-line versions of those algorithms.

References

Page 1

	Year	Citations

Page 1