EngineeringBusiness IntelligenceBusiness AnalyticsDatabase SystemData ScienceDatabase ProcessingManagementData IntegrationData ReductionInformation RedundancyMassive TablesData ManagementColumn DependencyData ModelingVery Large DatabaseComputer ScienceInformation ManagementDatabase TechnologyBusiness DataBig Data
Large amounts of business data are kept in tables of fixed-length records. Columns in such a table may be functionally dependent on one another, resulting in low overall information content. This paper shows how to exploit this source of information redundancy to compress table data. Experiments with a wide variety of massive tables including telecom data and stock quotes show that this technique compresses table data well, up to 48:1 or even 100:1 reduction in some cases.
17
The Pfam Protein Families Database
Alex Bateman · Nucleic Acids Research · 2002 · 14.2K citations · Full text
A Block-sorting Lossless Data Compression Algorithm
Michael T. Burrows, D. J. Wheeler · 1994 · 2.4K citations
The Pfam Protein Families Database
Alex Bateman · Nucleic Acids Research · 2000 · 1.3K citations · Full text