2013 · 17 citations · 6 references
Cluster ComputingEngineeringStorage SchemaAggregate FunctionInformation RetrievalData ScienceDatabase SupportManagementKeyvalue DatabaseData IntegrationData ManagementPractical QueriesComputer ScienceDistributed Query ProcessingValue IndexingQuery OptimizationQuery PerformanceEdge ComputingCloud ComputingApproximate Query AnsweringDistributed Data StoreBig Data
Open-source, BigTable-like distributed databases provide a scalable storage solution for data-intensive applications. The simple key-value storage schema provides fast record ingest and retrieval, nearly independent of the quantity of data stored. However, real applications must support non-trivial queries that require careful key design and value indexing. We study an Apache Accumulo-based big data system designed for a network situational awareness application. The application's storage schema and data retrieval requirements are analyzed. We then characterize the corresponding Accumulo performance bottlenecks. Queries are shown to be communication-bound and server-bound in different situations. Inefficiencies in the open-source communication stack and filesystem limit network and I/O performance, respectively. Additionally, in some situations, parallel clients can contend for server-side resources. Maximizing data retrieval rates for practical queries requires effective key design, indexing, and client parallelization.
6
Fay W. Chang, Sanjay Ghemawat, Wilson C. Hsieh et al. · ACM Transactions on Computer Systems · 2008 · 3.4K citations
A comparison of approaches to large-scale data analysis
Andrew Pavlo, Erik K. Paulson, Alexander Rasin et al. · 2009 · 1.1K citations
Swapnil Patil, Milo Polte, Kai Ren et al. · 2011 · 176 citations
Dynamic distributed dimensional data model (D4M) database and computation system
Jeremy Kepner, William Arcand, William Bergeron et al. · 2012 · 105 citations
Eric Anderson, Joseph Tucek · ACM SIGOPS Operating Systems Review · 2010 · 50 citations