IEEE Transactions on Computers · 2010 · 95 citations · 23 references
EngineeringGpu BenchmarkingHardware AlgorithmThroughput DriversComputer ArchitectureProcessor ArchitectureGpu ComputingHardware SecurityComputer DesignParallel ComputingGraphics ProcessorMemory Access RequirementsGraphics ProcessorsComputer EngineeringHeterogeneous SystemsComputer ScienceReconfigurable ArchitectureFpga DesignGpu ArchitectureHardware AccelerationCase StudyParallel ProgrammingPerformance Comparison
The study examines how analysis trends apply to newer and future technologies. The authors define a systematic approach to compare GPUs and reconfigurable logic based on three throughput drivers. They apply this approach to five algorithms differing in arithmetic complexity, memory access, and data dependence, using an NVIDIA GeForce 7900 GTX GPU and a Xilinx Virtex‑4 FPGA. The study finds that both the GPU and FPGA achieve two‑order‑of‑magnitude speedups over a general‑purpose processor for arithmetic‑intensive algorithms, that the FPGA is superior for algorithms with many regular memory accesses while the GPU is better for variable data reuse, and that for data‑dependent algorithms a customized FPGA datapath can outperform the GPU by up to eight times.
A systematic approach to the comparison of the graphics processor (GPU) and reconfigurable logic is defined in terms of three throughput drivers. The approach is applied to five case study algorithms, characterized by their arithmetic complexity, memory access requirements, and data dependence, and two target devices: the nVidia GeForce 7900 GTX GPU and a Xilinx Virtex-4 field programmable gate array (FPGA). Two orders of magnitude speedup, over a general-purpose processor, is observed for each device for arithmetic intensive algorithms. An FPGA is superior, over a GPU, for algorithms requiring large numbers of regular memory accesses, while the GPU is superior for algorithms with variable data reuse. In the presence of data dependence, the implementation of a customized data path in an FPGA exceeds GPU performance by up to eight times. The trends of the analysis to newer and future technologies are analyzed.
23
A Survey of General‐Purpose Computation on Graphics Hardware
John D. Owens, David Luebke, Naga K. Govindaraju et al. · Computer Graphics Forum · 2007 · 2K citations
John D. Owens, Michael Houston, David Luebke et al. · Proceedings of the IEEE · 2008 · 1.4K citations
Hardware Security, Cluster Computing, Computational Science +15
John D. Owens, Michael Houston, David Luebke et al. · Proceedings of the IEEE · 2008 · 1.1K citations
Hardware Security, Cluster Computing, Computational Science +15
Measuring the gap between FPGAs and ASICs
Ian Kuon, Jonathan Rose · 2006 · 318 citations