The K-mer File Format: a standardized and compact disk representation of sets of k-mers
- PMID: 35904548
- PMCID: PMC9477520
- DOI: 10.1093/bioinformatics/btac528
The K-mer File Format: a standardized and compact disk representation of sets of k-mers
Abstract
Summary: Bioinformatics applications increasingly rely on ad hoc disk storage of k-mer sets, e.g. for de Bruijn graphs or alignment indexes. Here, we introduce the K-mer File Format as a general lossless framework for storing and manipulating k-mer sets, realizing space savings of 3-5× compared to other formats, and bringing interoperability across tools.
Availability and implementation: Format specification, C++/Rust API, tools: https://github.com/Kmer-File-Format/.
Supplementary information: Supplementary data are available at Bioinformatics online.
© The Author(s) 2022. Published by Oxford University Press.
Figures

References
-
- Chikhi R. et al. (2021) Data structures to represent a set of k-long DNA sequences. ACM Comput. Surv., 54, 1–22.
-
- Cock P.J. et al. (2015) Sam/bam format v1. 5 extensions for de novo assemblies. BioRxiv, page 020024. 10.1101/020024. - DOI
Publication types
MeSH terms
Grants and funding
LinkOut - more resources
Full Text Sources
Miscellaneous