Pig latin
2008 · 1.7K citations · 12 references
Cluster ComputingEngineeringMap-reduceSoftware AnalysisData ScienceDatabase SupportManagementSql StyleData IntegrationParallel ComputingData ManagementData ModelingCustom User CodeComputer ScienceDatabase TechnologyData-intensive ComputingAd-hoc AnalysisParallel ProgrammingMassive Data ProcessingBig Data
There is a growing need for ad-hoc analysis of extremely large data sets, especially at internet companies where innovation critically depends on being able to analyze terabytes of data collected every day. Parallel database products, e.g., Teradata, offer a solution, but are usually prohibitively expensive at this scale. Besides, many of the people who analyze this data are entrenched procedural programmers, who find the declarative, SQL style to be unnatural. The success of the more procedural map-reduce programming model, and its associated scalable implementations on commodity hardware, is evidence of the above. However, the map-reduce paradigm is too low-level and rigid, and leads to a great deal of custom user code that is hard to maintain, and reuse.
12
Sanjay Ghemawat · Communications of the ACM · 2008
18.4K citations
Giuseppe DeCandia, Deniz Hastorun, Madan Jampani et al. · 2007
3.4K citations
Fay W. Chang, Sanjay Ghemawat, Wilson C. Hsieh et al. · ACM Transactions on Computer Systems · 2008
3.4K citations
Michael Isard, Mihai Budiu, Yuan Yu et al. · 2007
2.4K citations
Data Cube: A Relational Aggregation Operator Generalizing Group-By, Cross-Tab, and Sub-Totals
Jim Gray, Surajit Chaudhuri, Adam Bosworth et al. · Data Mining and Knowledge Discovery · 1997
2.1K citations