IEEE Transactions on Computers · 2016 · 53 citations · 34 references
Hardware ModelingEngineeringEnergy EfficiencyPower Optimization (Eda)Computer ArchitecturePower OptimizationDennard ScalingHardware SystemsProcessor ArchitectureHigh-performance ArchitectureModeling And SimulationParallel ComputingManycore ProcessorElectrical EngineeringAnalytical ModelsComputer EngineeringComputer ScienceSpecific ApplicationHardware AccelerationAnalytical Processor PerformanceMany-core ArchitectureMultiprocessor SystemParallel Programming
Optimizing processors for (a) specific application(s) can substantially improve energy-efficiency. With the end of Dennard scaling, and the corresponding reduction in energy-efficiency gains from technology scaling, such approaches may become increasingly important. However, designing application-specific processors requires fast design space exploration tools to optimize for the targeted application(s). Analytical models can be a good fit for such design space exploration as they provide fast performance and power estimates and insight into the interaction between an application's characteristics and the micro-architecture of a processor. Unfortunately, prior analytical models for superscalar out-of-order processors require micro-architecture dependent inputs, such as cache miss rates, branch miss rates and memory-level parallelism. This requires profiling the applications for each cache and branch predictor configuration of interest, which is far more time-consuming than evaluating the analytical performance models. In this work we present a <i>micro-architecture independent</i> profiler and associated analytical models that allow us to produce performance <i>and</i> power estimates across a large superscalar out-of-order processor design space almost instantaneously. We show that using a micro-architecture independent profile leads to a speedup of 300 <inline-formula><tex-math notation="LaTeX">$\times$</tex-math></inline-formula> compared to detailed simulation for our evaluated design space. Over a large design space, the model has a 9.3 percent average error for performance and a 4.3 percent average error for power, compared to detailed cycle-level simulation. The model is able to accurately determine the optimal processor configuration for different applications under power or performance constraints, and provides insight into performance through cycle stacks.
34
Design of ion-implanted MOSFET's with very small physical dimensions
R.H. Dennard, F.H. Gaensslen, Hwa-Nien Yu et al. · IEEE Journal of Solid-State Circuits · 1974 · 3.4K citations
Chi-Keung Luk, Robert Cohn, Robert Muth et al. · ACM SIGPLAN Notices · 2005 · 3.2K citations
Engineering, Computer Architecture, Software Engineering +16
David Brooks, Vivek Tiwari, Margaret Martonosi · 2000 · 2.6K citations
Engineering, Energy Efficiency, Power Optimization (Eda) +16
Sheng Li, Jung Ho Ahn, Richard Strong et al. · 2009 · 2.3K citations
Hardware Security, Electrical Engineering, Hardware Modeling +13
Chi-Keung Luk, Robert Cohn, Robert Muth et al. · 2005 · 2.3K citations
Engineering, Computer Architecture, Software Engineering +16