Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2015 Jun;33(6):623-30.
doi: 10.1038/nbt.3238. Epub 2015 May 25.

Assembling large genomes with single-molecule sequencing and locality-sensitive hashing

Affiliations

Assembling large genomes with single-molecule sequencing and locality-sensitive hashing

Konstantin Berlin et al. Nat Biotechnol. 2015 Jun.

Erratum in

Abstract

Long-read, single-molecule real-time (SMRT) sequencing is routinely used to finish microbial genomes, but available assembly methods have not scaled well to larger genomes. We introduce the MinHash Alignment Process (MHAP) for overlapping noisy, long reads using probabilistic, locality-sensitive hashing. Integrating MHAP with the Celera Assembler enabled reference-grade de novo assemblies of Saccharomyces cerevisiae, Arabidopsis thaliana, Drosophila melanogaster and a human hydatidiform mole cell line (CHM1) from SMRT sequencing. The resulting assemblies are highly continuous, include fully resolved chromosome arms and close persistent gaps in these reference genomes. Our assembly of D. melanogaster revealed previously unknown heterochromatic and telomeric transition sequences, and we assembled low-complexity sequences from CHM1 that fill gaps in the human GRCh38 reference. Using MHAP and the Celera Assembler, single-molecule sequencing can produce de novo near-complete eukaryotic assemblies that are 99.99% accurate when compared with available reference genomes.

PubMed Disclaimer

References

    1. PLoS Comput Biol. 2014 Dec 04;10(12):e1003998 - PubMed
    1. Genome Res. 2010 Sep;20(9):1165-73 - PubMed
    1. Genome Res. 2012 Nov;22(11):2270-7 - PubMed
    1. BMC Genomics. 2012 Sep 13;13:475 - PubMed
    1. Open Biol. 2012 Aug;2(8):120093 - PubMed

Publication types

Substances

LinkOut - more resources