2013 · 37 citations · 33 references
Distributed File SystemCluster ComputingStorage PerformanceEngineeringComputer ArchitectureData DeduplicationBlock LocalityData ConsistencyData ScienceLocality InformationManagementData IntegrationParallel ComputingData ManagementData BlocksComputer EngineeringData PrivacyCachingChunk IndexComputer ScienceCloud ComputingParallel ProgrammingBig Data
Data deduplication systems discover and remove redundancies between data blocks by splitting the data stream into chunks and comparing a hash of each chunk with all previously stored hashes. Storing the corresponding chunk index on hard disks immediately limits the achievable throughput, as these devices are unable to support the high number of random IOs induced by this index. Several approaches to overcome this chunk lookup disk bottleneck have been proposed. Often, the approaches try to capture the locality information of a backup run and use this in the next backup run to predict future chunk requests. However, often this locality is only captured by a surrogate, e.g., the order of the chunks in containers. [37]. Furthermore, some approaches degenerate slowly when the systems operate over months and years because the locality information becomes outdated.
33
Space/time trade-offs in hash coding with allowable errors
Burton H. Bloom · Communications of the ACM · 1970 · 7.4K citations · Full text
A low-bandwidth network file system
Athicha Muthitacharoen, Benjie Chen, David Mazières · 2001 · 782 citations
Hardware Security, Cluster Computing, Distributed File System +14
Venti: A New Approach to Archival Storage
Sean Quinlan, Sean Dorward · 2002 · 726 citations
Avoiding the disk bottleneck in the data domain deduplication file system
Benjamin Zhu, Li Kai, Hugo Patterson · 2008 · 692 citations
Sparse indexing: large scale, inline deduplication using sampling and locality
Mark Lillibridge, Kave Eshghi, Deepavali Bhagwat et al. · 2009 · 382 citations