Proceedings of the IEEE · 2018 · 219 citations · 50 references
EngineeringMachine LearningDeep Learning HardwareDeep Learning SpaceHardware AccelerationSparse Neural NetworkAnalog DesignNeurocomputersComputer EngineeringComputer ArchitectureParallel ProgrammingComputer ScienceBrain-like ComputingParallel ComputingDeep LearningDeep Learning WorkIn-memory ComputingGpu Computing
GPUs, originally designed for gaming and 3‑D rendering, are well suited to accelerate deep learning due to their parallelizable mathematical structure, but current nonvolatile memory materials limit the potential of in‑memory compute paradigms. The study aims to improve deep learning compute efficiency by exploiting its tolerance for randomness and approximation, and by exploring noisy analog computing with nonvolatile memory arrays for constant‑time matrix operations. The approach trades numerical precision for accuracy to boost compute efficiency and revisits analog computing using noisy nonvolatile memory arrays for constant‑time matrix operations. Analysis and design guidelines reveal that materials must be reengineered for deep learning, diverging significantly from those used in conventional memory applications.
Initially developed for gaming and 3-D rendering, graphics processing units (GPUs) were recognized to be a good fit to accelerate deep learning training. Its simple mathematical structure can easily be parallelized and can therefore take advantage of GPUs in a natural way. Further progress in compute efficiency for deep learning training can be made by exploiting the more random and approximate nature of deep learning work flows. In the digital space that means to trade off numerical precision for accuracy at the benefit of compute efficiency. It also opens the possibility to revisit analog computing, which is intrinsically noisy, to execute the matrix operations for deep learning in constant time on arrays of nonvolatile memories. To take full advantage of this in-memory compute paradigm, current nonvolatile memory materials are of limited use. A detailed analysis and design guidelines how these materials need to be reengineered for optimal performance in the deep learning space shows a strong deviation from the materials used in memory applications.
50
Sepp Hochreiter, Jürgen Schmidhuber · Neural Computation · 1997 · 93.8K citations
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio et al. · Proceedings of the IEEE · 1998 · 56.5K citations · Full text
Engineering, Machine Learning, Multilayer Neural Networks +17
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams · Nature · 1986 · 29.7K citations
In-Datacenter Performance Analysis of a Tensor Processing Unit
Norman P. Jouppi, Cliff Young, Nishant Patil et al. · 2017 · 4.3K citations