2016 · 115 citations · 48 references
EngineeringMemory DesignComputer ArchitectureData BuffersHardware SystemsMulti-channel Memory Architecture3D MemoryNda ArchitectureHigh-performance ArchitectureComputing SystemsMemory DevicesParallel ComputingComputer EngineeringLarge Memory SystemsComputer ScienceMicroelectronicsMemory ArchitectureNear-dram AcceleratorsHardware AccelerationHigh Bandwidth MemoryMany-core ArchitectureIn-memory Computing
Memory bandwidth limits system performance, and existing near‑DRAM acceleration designs rely on costly high‑bandwidth memory with limited capacity, hindering scalability for server‑grade systems. The paper proposes Chameleon, a near‑DRAM accelerator that can be integrated into large LRDIMM memory systems without 3D/2.5D stacking. The authors explore three microarchitectures that mitigate the constraints imposed by using LRDIMM for near‑DRAM acceleration. Experiments show that a Chameleon‑based system delivers 2.13× higher geo‑mean performance while reducing geo‑mean data‑transfer energy by 34 % compared to embedding the same accelerator logic in the processor.
The performance of computer systems is often limited by the bandwidth of their memory channels, but further increasing the bandwidth is challenging under the stringent pin and power constraints of packages. To further increase performance under these constraints, various near-DRAM acceleration (NDA) architectures, which tightly integrate accelerators with DRAM devices using 3D/2.5D-stacking technology, have been proposed. However, they have not prevailed yet because they often rely on expensive HBM/HMC-like DRAM devices which also suffer from limited capacity, whereas the scalability of memory capacity is critical for some computing segments such as servers. In this paper, we first demonstrate that data buffers in a load-reduced DIMM (LRDIMM), which is originally developed to support large memory systems for servers, are supreme places to integrate near-DRAM accelerators. Second, we propose Chameleon, an NDA architecture that can be realized without relying on 3D/2.5D-stacking technology and seamlessly integrated with large memory systems for servers. Third, we explore three microarchitectures that abate constraints imposed by taking LRDIMM architecture for NDA. Our experiment demonstrates that a Chameleon-based system can offer 2.13 χ higher geo-mean performance while consuming 34% lower geo-mean data transfer energy than a system that integrates the same accelerator logic within the processor.
48
Rodinia: A benchmark suite for heterogeneous computing
Shuai Che, Michael Boyer, Jiayuan Meng et al. · 2009 · 3K citations
The SPLASH-2 programs: characterization and methodological considerations
Sandi Woo, Moriyoshi Ohara, Evan Torrie et al. · 2002 · 1.6K citations
R. Ho, Ken Mai, Mark Horowitz · Proceedings of the IEEE · 2001 · 1.4K citations
Electrical Engineering, Physical Design (Electronics), Engineering +15