2014 · 37 citations · 8 references
Cluster ComputingEngineeringDatabase ScalabilityMap-reduceCluster TechnologyData ScienceData IntegrationParallel ComputingSimple QueriesData ManagementHigh-performance Data AnalyticsComputer EngineeringMysql ClusterComputer ScienceDistributed Query ProcessingData-intensive ComputingFamous Clustered DatabasePerformance ScalabilityApache PigParallel Performance EvaluationCloud ComputingParallel ProgrammingMassive Data ProcessingBig Data
MySQL Cluster is a famous clustered database that is used to store and manipulate data. The problem with MySQL Cluster is that as the data grows larger, the time required to process the data increases and additional resources may be needed. With Hadoop and Hive and Pig, processing time can be faster than MySQL Cluster. In this paper, three data testers with the same data model will run simple queries and to find out at how many rows Hive or Pig is faster than MySQL Cluster. The data model taken from GroupLens Research Project [12] showed a result that Hive is the most appropriate for this data model in a low-cost hardware environment.
8
The Hadoop Distributed File System
Konstantin V. Shvachko, Hairong Kuang, Sanjay Radia et al. · 2010 · 4.8K citations
Christopher Olston, Benjamin Reed, Utkarsh Srivastava et al. · 2008 · 1.7K citations
Ashish Thusoo, Joydeep Sen Sarma, Namit Jain et al. · Proceedings of the VLDB Endowment · 2009 · 1.5K citations
Hive - a petabyte scale data warehouse using Hadoop
Ashish Thusoo, Joydeep Sen Sarma, Namit Jain et al. · 2010 · 920 citations
Parallel data processing with MapReduce
Kyong-Ha Lee, Yoon-Joon Lee, Hyun‐Sik Choi et al. · ACM SIGMOD Record · 2012 · 555 citations