Building the sequence map of the human pan-genome
- PMID: 19997067
- DOI: 10.1038/nbt.1596
Building the sequence map of the human pan-genome
Abstract
Here we integrate the de novo assembly of an Asian and an African genome with the NCBI reference human genome, as a step toward constructing the human pan-genome. We identified approximately 5 Mb of novel sequences not present in the reference genome in each of these assemblies. Most novel sequences are individual or population specific, as revealed by their comparison to all available human DNA sequence and by PCR validation using the human genome diversity cell line panel. We found novel sequences present in patterns consistent with known human migration paths. Cross-species conservation analysis of predicted genes indicated that the novel sequences contain potentially functional coding regions. We estimate that a complete human pan-genome would contain approximately 19-40 Mb of novel sequence not present in the extant reference genome. The extensive amount of novel sequence contributing to the genetic variation of the pan-genome indicates the importance of using complete genome sequencing and de novo assembly.
Similar articles
-
NSIT: novel sequence identification tool.PLoS One. 2014 Sep 29;9(9):e108011. doi: 10.1371/journal.pone.0108011. eCollection 2014. PLoS One. 2014. PMID: 25264906 Free PMC article.
-
[Analysis, identification and correction of some errors of model refseqs appeared in NCBI Human Gene Database by in silico cloning and experimental verification of novel human genes].Yi Chuan Xue Bao. 2004 May;31(5):431-43. Yi Chuan Xue Bao. 2004. PMID: 15478601 Chinese.
-
Building a sequence map of the pig pan-genome from multiple de novo assemblies and Hi-C data.Sci China Life Sci. 2020 May;63(5):750-763. doi: 10.1007/s11427-019-9551-7. Epub 2019 Jul 8. Sci China Life Sci. 2020. PMID: 31290097
-
Whole genome sequencing.Methods Mol Biol. 2010;628:215-26. doi: 10.1007/978-1-60327-367-1_12. Methods Mol Biol. 2010. PMID: 20238084 Review.
-
State of the art de novo assembly of human genomes from massively parallel sequencing data.Hum Genomics. 2010 Apr;4(4):271-7. doi: 10.1186/1479-7364-4-4-271. Hum Genomics. 2010. PMID: 20511140 Free PMC article. Review.
Cited by
-
When Livestock Genomes Meet Third-Generation Sequencing Technology: From Opportunities to Applications.Genes (Basel). 2024 Feb 15;15(2):245. doi: 10.3390/genes15020245. Genes (Basel). 2024. PMID: 38397234 Free PMC article. Review.
-
The Need for a Human Pangenome Reference Sequence.Annu Rev Genomics Hum Genet. 2021 Aug 31;22:81-102. doi: 10.1146/annurev-genom-120120-081921. Epub 2021 Apr 30. Annu Rev Genomics Hum Genet. 2021. PMID: 33929893 Free PMC article. Review.
-
Maize (Zea mays L.) genome diversity as revealed by RNA-sequencing.PLoS One. 2012;7(3):e33071. doi: 10.1371/journal.pone.0033071. Epub 2012 Mar 16. PLoS One. 2012. PMID: 22438891 Free PMC article.
-
Population-genetic properties of differentiated human copy-number polymorphisms.Am J Hum Genet. 2011 Mar 11;88(3):317-32. doi: 10.1016/j.ajhg.2011.02.004. Am J Hum Genet. 2011. PMID: 21397061 Free PMC article.
-
A lizard is never late: squamate genomics as a recent catalyst for understanding sex chromosome and microchromosome evolution.bioRxiv [Preprint]. 2023 Mar 31:2023.01.20.524006. doi: 10.1101/2023.01.20.524006. bioRxiv. 2023. Update in: J Hered. 2023 Aug 23;114(5):445-458. doi: 10.1093/jhered/esad023. PMID: 37034614 Free PMC article. Updated. Preprint.
References
Publication types
MeSH terms
LinkOut - more resources
Full Text Sources
Molecular Biology Databases
Miscellaneous