2004 · 85 citations · 11 references
Hardware SecurityArray ComputingEngineeringHardware AccelerationEntire FpgaHardware AlgorithmComputer EngineeringComputer ArchitectureLinear Array ArchitectureParallel ProgrammingComputer ScienceReconfigurable ArchitectureFundamental KernelParallel ComputingModular AlgorithmsFpga Design
Summary form only given. The abundant hardware resources on current FPGAs provide new opportunities to improve the performance of hardware implementations of scientific computations. We propose two FPGA-based algorithms for floating-point matrix multiplication, a fundamental kernel in a number of scientific applications. We analyze the design tradeoffs in implementing this kernel on FPGAs. Our algorithms employ a linear array architecture with a small control logic. This architecture effectively utilizes the hardware resources on the entire FPGA and reduces the routing complexity. The processing elements (PEs) used in our algorithms are modular so that floating-point units can be easily embedded into them. In our designs, the floating-point units are optimized to maximize the number of PEs integrated on the FPGA as well as the clock speed. Experimental results show that our algorithms achieve high clock speeds and provide good scalability. Our algorithms achieve superior sustained floating-point performance compared with existing FPGA-based implementations and state-of-the-art processors.
11
THE INSTITUTEOF ELECTRICAL AND ELECTRONICS ENGINEERS,INC.
Hendley Blackmon, E. L. Harder, S. Saito et al. · 1966 · 1.4K citations
Sendthemanuscript Toe.k.gannett, Electrical Engineering, Engineering +10
Jiawei Hong, H. T. Kung · 1981 · 448 citations
Theory Of Computing, Computational Complexity Theory, Engineering +13
Analysis of high-performance floating-point arithmetic on FPGAs
Gokul Govindu, Ling Zhuo, Seonil Choi et al. · 2004 · 136 citations
A re-evaluation of the practicality of floating-point operations on FPGAs
W.B. Ligon, Scott McMillan, G. Monn et al. · 2002 · 111 citations