2018 · 42 citations · 13 references
EngineeringMachine LearningComputer ArchitectureChipmunk ArchitectureHardware SystemsRecurrent Neural NetworkSpeech RecognitionScalable 0.9Computing SystemsRecurrent Neural NetworksReal-time LanguageMw AcceleratorMultiple Chipmunk EnginesComputer EngineeringComputer ScienceHardware AccelerationDomain-specific AcceleratorSpeech ProcessingBrain-like ComputingTechnology
Recurrent neural networks (RNNs) are state-of-the-art in voice awareness/understanding and speech recognition. On-device computation of RNNs on low-power mobile and wearable devices would be key to applications such as zero-latency voice-based human-machine interfaces. Here we present CHIPMUNK, a small (<;1 mm <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> ) hardware accelerator for Long-Short Term Memory RNNs in UMC 65 nm technology capable to operate at a measured peak efficiency up to 3.08Gop/s/mW at 1.24 mW peak power. To implement big RNN models without incurring in huge memory transfer overhead, multiple CHIPMUNK engines can cooperate to form a single systolic array. In this way, the Chipmunk architecture in a 75 tiles configuration can achieve real-time phoneme extraction on a demanding RNN topology proposed in [1], consuming less than 13 mW of average power.
13
Sepp Hochreiter, Jürgen Schmidhuber · Neural Computation · 1997 · 93.8K citations
In-Datacenter Performance Analysis of a Tensor Processing Unit
Norman P. Jouppi, Cliff Young, Nishant Patil et al. · 2017 · 4.3K citations