2005 · 92 citations · 34 references
Conventional data integration relies on a top‑down schema design that works for static enterprise data but clashes with the bottom‑up model of scientific data sharing, where new data must be rapidly published, revised, and reconciled by multiple parties. We aim to enable bottom‑up collaborative data sharing among independent researchers with differing schemas and goals, without requiring global agreement. ORCHESTRA lets each group independently curate and extend its data, then compare and reconcile changes without mandatory consensus, focusing on managing disagreement among multiple data representations.
Conventional data integration techniques employ a “top-down” design philosophy, starting by assessing requirements and defining a global schema, and then mapping data sources to that schema. This works well if the problem domain is well-understood and relatively static, as with enterprise data. However, it is fundamentally mismatched with the “bottom-up”model of scientific data sharing, in which new data needs to be rapidly developed, published, and then assessed, filtered, and revised by others. We address the need for bottom-up collaborative data sharing, in which independent researchers or groups with different goals, schemas, and data can share information in the absence of global agreement. Each group independently curates, revises, and extends its data; eventually the groups compare and reconcile their changes, but they are not required to agree. This paper describes our initial design and prototype of the ORCHESTRA system, which focuses on managing disagreement among multiple data representations and instances. Our work represents an important evolution of the concepts of peer-to-peer data sharing [23], which considers revision, disagreement, authority, and intermittent participation.
34
Ion Stoica, Robert Morris, David R. Karger et al. · 2001 · 9.6K citations
A scalable content-addressable network
Sylvia Ratnasamy, Paul Francis, Mark Handley et al. · 2001 · 6.4K citations · Full text
A survey of approaches to automatic schema matching
Erhard Rahm, Philip A. Bernstein · The VLDB Journal · 2001 · 3.3K citations
Epidemic algorithms for replicated database maintenance
Alan Demers, Dan Greene, Carl Hauser et al. · 1987 · 1.6K citations · Full text