2013 · 38 citations · 18 references
Effective UseEngineeringComputer ArchitectureEmbedded SystemsProcessor ArchitectureHardware SystemsHardware SecurityHigh-performance ArchitectureParallel ComputingCompilersInstruction-level ParallelismNon-temporal Store GenerationParallelizing CompilerXeon PhiComputer EngineeringComputer ScienceParallel ApplicationsOperating SystemsProgram AnalysisSpecial Store InstructionsCompiler-based Data PrefetchingParallel ProgrammingReal-time SystemsSystem Software
The Intel® Xeon Phi™ coprocessor has software prefetching instructions to hide memory latencies and special store instructions to save bandwidth on streaming non-temporal store operations. In this work, we provide details on compiler-based generation of these instructions and evaluate their impact on the performance of the Intel® Xeon Phi™ coprocessor using a wide range of parallel applications with different characteristics. Our results show that the Intel® Composer XE 2013 compiler can make effective use of these mechanisms to achieve significant performance improvements.
18
Rodinia: A benchmark suite for heterogeneous computing
Shuai Che, Michael Boyer, Jiayuan Meng et al. · 2009 · 3K citations
The NAS parallel benchmarks---summary and preliminary results
David H. Bailey, Robert Schreiber, Horst D. Simon et al. · 1991 · 567 citations
David Callahan, Ken Kennedy, Allan Porterfield · 1991 · 486 citations
An effective on-chip preloading scheme to reduce data access penalty
Jean-Loup Baer, Tien-Fu Chen · 1991 · 459 citations · Full text