2011 · 61 citations · 7 references
Cluster ComputingLossy CompressionEngineeringReal-time Data CompressionComputer ArchitectureGpu ComputingHardware SecurityHigh-performance ArchitectureNumeric SimulationsParallel ComputingLossless CompressionFloating-point Data CompressionComputer EngineeringNetwork CardsComputer ScienceData CompressionGpu ClusterGpu ArchitectureHardware AccelerationEdge ComputingCloud ComputingParallel Programming
Numeric simulations often generate large amounts of data that need to be stored or sent to other compute nodes. This paper investigates whether GPUs are powerful enough to make real-time data compression and decompression possible in such environments, that is, whether they can operate at the 32- or 40-Gb/s throughput of emerging network cards. The fastest parallel CPU-based floating-point data compression algorithm operates below 20 Gb/s on eight Xeon cores, which is significantly slower than the network speed and thus insufficient for compression to be practical in high-end networks. As a remedy, we have created the highly parallel GFC compression algorithm for double-precision floating-point data. This algorithm is specifically designed for GPUs. It compresses at a minimum of 75 Gb/s, decompresses at 90 Gb/s and above, and can therefore improve internode communication throughput on current and upcoming networks by fully saturating the interconnection links with compressed data.
7
NVIDIA Tesla: A Unified Graphics and Computing Architecture
Erik Lindholm, John Nickolls, Stuart F. Oberman et al. · IEEE Micro · 2008 · 1.3K citations
Parallel Prefix Sum (Scan) with CUDA
Mark Harris · 2011 · 541 citations