2010 · 11 citations · 11 references
Cluster ComputingEngineeringMap-reduceData ScienceData MiningData-intensive PlatformManagementData IntegrationParallel ComputingData ManagementHigh-performance Data AnalyticsKnowledge DiscoveryInternet-scale IndexingComputer ScienceData-intensive ComputingParallel Bulk InsertionPopular Mapreduce FrameworkData ProcessingParallel ProgrammingParallel ApproachData-level ParallelismMassive Data ProcessingBig Data
Modern data analytics applications, e.g. Internet-scale indexing, system trace analysis, recommender engines to name a few, operate on massive amounts of data and call for a parallel approach to data processing. In this work, we focus on the popular MapReduce framework to carry out such tasks and identify bulk data insert operations as a critical preliminary step to achieve reduced processing times, especially when new data is generated and processed at regular time intervals.
11
Sanjay Ghemawat · Communications of the ACM · 2008 · 18.4K citations · Full text
Benchmarking cloud serving systems with YCSB
Brian F. Cooper, Adam Silberstein, Erwin Tam et al. · 2010 · 3.6K citations
Giuseppe DeCandia, Deniz Hastorun, Madan Jampani et al. · 2007 · 3.4K citations
Fay W. Chang, Sanjay Ghemawat, Wilson C. Hsieh et al. · ACM Transactions on Computer Systems · 2008 · 3.4K citations
Avinash Lakshman, Prashant Malik · ACM SIGOPS Operating Systems Review · 2010 · 2.6K citations