2015 · 133 citations · 65 references
Approximate execution benefits many accelerator‑eligible applications, and prior work demonstrates neural acceleration as a viable approach. The study investigates neural acceleration on commercially available programmable SoCs. SNNAP is a flexible FPGA‑based neural accelerator that uses a compiler workflow to configure the neural network’s topology and weights, avoiding direct FPGA logic programming. SNNAP delivers up to 38× speedup and 28× energy savings with <10% quality loss, requires no hardware expertise, and achieves comparable or better performance than commercial HLS tools on most benchmarks.
Many applications that can take advantage of accelerators are amenable to approximate execution. Past work has shown that neural acceleration is a viable way to accelerate approximate code. In light of the growing availability of on-chip field-programmable gate arrays (FPGAs), this paper explores neural acceleration on off-the-shelf programmable SoCs. We describe the design and implementation of SNNAP, a flexible FPGA-based neural accelerator for approximate programs. SNNAP is designed to work with a compiler workflow that configures the neural network's topology and weights instead of the programmable logic of the FPGA itself. This approach enables effective use of neural acceleration in commercially available devices and accelerates different applications without costly FPGA reconfigurations. No hardware expertise is required to accelerate software with SNNAP, so the effort required can be substantially lower than custom hardware design for an FPGA fabric and possibly even lower than current "C-to-gates" high-level synthesis (HLS) tools. Our measurements on a Xilinx Zynq FPGA show that SNNAP yields a geometric mean of 3.8× speedup (as high as 38.1×) and 2.8× energy savings (as high as 28 x) with less than 10% quality loss across all applications but one. We also compare SNNAP with designs generated by commercial HLS tools and show that SNNAP has similar performance overall, with better resource-normalized throughput on 4 out of 7 benchmarks.
65
UCI Machine Learning Repository
Arthur Asuncion · Medical Entomology and Zoology · 2007 · 24.3K citations
Christian Bienia, Sanjeev Kumar, Jaswinder Pal Singh et al. · 2008 · 3.4K citations
Synthesis and optimization of digital circuits
Choice Reviews Online · 1994 · 2.4K citations
Dark silicon and the end of multicore scaling
Hadi Esmaeilzadeh, Emily Blem, Renée St. Amant et al. · 2011 · 1.5K citations
Heterogeneous Computing, Engineering, Computer Architecture +16
Tianshi Chen, Zidong Du, Ninghui Sun et al. · 2014 · 1.3K citations
Deep Neural Networks, Machine-learning Algorithms, Machine Learning +15