Concepedia

Scientific data management in the coming decade

Jim Gray, David T. Liu, M. A. Nieto‐Santisteban, Alexander S. Szalay, David J. DeWitt, Gerd Heber

ACM SIGMOD Record · 2005 · 483 citations · 6 references

Concepts

Abstract

Scientific instruments and computer simulations are creating vast data stores that require new scientific methods to analyze and organize the data. Data volumes are approximately doubling each year. Since these new instruments have extraordinary precision, the data quality is also rapidly improving. Analyzing this data to find the subtle effects missed by previous studies requires algorithms that can simultaneously deal with huge datasets and that can find very subtle effects --- finding both needles in the haystack and finding very small haystacks that were undetected in previous measurements.

References

6

Globus: a Metacomputing Infrastructure Toolkit

Ian Foster, Carl Kesselman · The International Journal of Supercomputer Applications and High Performance Computing · 1997

+21

3.1K citations

MapReduce: Simplified Data Processing on Large Cluster

INTERNATIONAL JOURNAL OF RESEARCH AND ENGINEERING · 2018

3K citations

2.3K citations

Parallel database systems

David J. DeWitt, Jim Gray · Communications of the ACM · 1992

1.3K citations

When Database Systems Meet the Grid

M. A. Nieto‐Santisteban, Alexander S. Szalay, Aniruddha R. Thakar et al. · ArXiv.org · 2005

+12

47 citations