Data compression for sequencing data
- PMID: 24252160
- PMCID: PMC3868316
- DOI: 10.1186/1748-7188-8-25
Data compression for sequencing data
Abstract
: Post-Sanger sequencing methods produce tons of data, and there is a general agreement that the challenge to store and process them must be addressed with data compression. In this review we first answer the question "why compression" in a quantitative manner. Then we also answer the questions "what" and "how", by sketching the fundamental compression ideas, describing the main sequencing data types and formats, and comparing the specialized compression algorithms and tools. Finally, we go back to the question "why compression" and give other, perhaps surprising answers, demonstrating the pervasiveness of data compression techniques in computational biology.
Figures
References
-
- Metzker ML. Sequencing technologies–the next generation. Nat Rev Genet. 2010;11:31–46. - PubMed
-
- Kahn SD. On the future of genomic data. Science. 2011;331:728–729. - PubMed
-
- Roberts JP. Million veterans sequenced. Nat Biotechnol. 2013;31(6):470.
-
- National Human Genome Research Institute, DNA Sequencing Costs. [ http://www.genome.gov/sequencingcosts/] (accessed February 14, 2013)
LinkOut - more resources
Full Text Sources
Other Literature Sources
