2018 · 33 citations · 17 references
Cluster ComputingEngineeringComputer ArchitectureMap-reduceData ScienceReference-distance EvictionManagementData IntegrationParallel ComputingData ManagementWeb CacheHigh-performance Data AnalyticsMemory Cache UsageComputer EngineeringCachingComputer ScienceApplication RuntimeData-intensive ComputingExternal-memory AlgorithmParallel ProgrammingDag StructureIn-memory DatabaseBig Data
Optimizing memory cache usage is vital for performance of in-memory data-parallel frameworks such as Spark. Current data-analytic frameworks utilize the popular Least Recently Used (LRU) policy, which does not take advantage of data dependency information available in the application's directed acyclic graph (DAG). Recent research in dependency-aware caching, notably MemTune and Least Reference Count (LRC), have made important improvements to close this gap. But they do not fully leverage the DAG structure, which imparts information such as the time-spatial distribution of data references across the workflow, to further improve cache hit ratio and application runtime.
17
The Hadoop Distributed File System
Konstantin V. Shvachko, Hairong Kuang, Sanjay Radia et al. · 2010 · 4.8K citations
Evaluation techniques for storage hierarchies
R. L. Mattson, J. Gecsei, Donald R. Slutz et al. · IBM Systems Journal · 1970 · 1.3K citations
Haoyuan Li, Ali Ghodsi, Matei Zaharia et al. · 2014 · 297 citations
Andrew D. Ferguson, Peter Bodík, Srikanth Kandula et al. · 2012 · 292 citations
Job Scheduler, Cluster Computing, Data Processing Frameworks +15