arXiv (Cornell University) · 2011 · 47 citations · 22 references
Open access
We present a system and a set of techniques for learning linear predictors with convex losses on terascale datasets, with trillions of features,1 billions of training examples and millions of parameters in an hour using a cluster of 1000 machines. One of the core techniques used is a new communication infrastructure—often re-ferred to as AllReduce—implemented for compatibility with MapReduce clusters. The communication infrastructure appears broadly reusable for many other tasks. We also show the effectiveness of a hybrid online-batch approach for optimization in distributed settings. 1
22
Sanjay Ghemawat · Communications of the ACM · 2008 · 18.4K citations · Full text
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
John C. Duchi, Elad Hazan, Yoram Singer · 2010 · 8.6K citations
Resilient distributed datasets: a fault-tolerant abstraction for in-memory cluster computing
Matei Zaharia, Mosharaf Chowdhury, Tathagata Das et al. · 2012 · 3.6K citations
Large Scale Distributed Deep Networks
Jay B. Dean, Greg S. Corrado, Rajat Monga et al. · 2012 · 2.9K citations