2007 · 60 citations · 40 references
We present STAR, a self-tuning algorithm that adaptively sets numeric precision constraints to accurately and efficiently answer continuous aggregate queries over distributed data streams. Adaptivity and approximation are essential for both robustness to varying workload characteristics and for scalability to large systems. In contrast to previous studies, we treat the problem as a workload-aware optimization problem whose goal is to minimize the total communication load for a multi-level aggregation tree under a fixed error budget. STAR’s hierarchical algorithm takes into account the update rate and variance in the input data distribution in a principled manner to compute an optimal error distribution, and it performs cost-benefit throttling to direct error slack to where it yields the largest benefits. Our prototype implementation of STAR in a large-scale monitoring system provides (1) a new distribution mechanism that enables selftuning error distribution and (2) an optimization to reduce communication overhead in a practical setting by carefully distributing the initial, default error budgets. Through extensive simulations and experiments on a real network monitoring implementation, we show that STAR achieves significant performance benefits compared to existing approaches while still providing high accuracy and incurring low overheads.
40
Ion Stoica, Robert Morris, David R. Karger et al. · 2001 · 9.6K citations
A scalable content-addressable network
Sylvia Ratnasamy, Paul Francis, Mark Handley et al. · 2001 · 6.4K citations · Full text
Samuel Madden, Michael J. Franklin, Joseph M. Hellerstein et al. · ACM SIGOPS Operating Systems Review · 2002 · 2.7K citations
Models and issues in data stream systems
Brian Babcock, Shivnath Babu, Mayur Datar et al. · 2002 · 2.5K citations