NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Kim D Pruitt¹, Tatiana Tatusova, Donna R Maglott

Affiliations

Affiliation

¹ National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Rm 6An.12J, 45 Center Drive, Bethesda, MD 20892-6510, USA. pruitt@ncbi.nlm.nih.gov

PMID: 17130148
PMCID: PMC1716718
DOI: 10.1093/nar/gkl842

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Kim D Pruitt et al. Nucleic Acids Res. 2007 Jan.

. 2007 Jan;35(Database issue):D61-5.

doi: 10.1093/nar/gkl842. Epub 2006 Nov 27.

Authors

Kim D Pruitt¹, Tatiana Tatusova, Donna R Maglott

Affiliation

¹ National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Rm 6An.12J, 45 Center Drive, Bethesda, MD 20892-6510, USA. pruitt@ncbi.nlm.nih.gov

PMID: 17130148
PMCID: PMC1716718
DOI: 10.1093/nar/gkl842

Abstract

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.

PubMed Disclaimer

Figures

**Figure 1**
Entrez query results include records from RefSeq and GenBank (nucleotide queries) or GenPept (protein queries). (A) Users who register for MyNCBI can log on to access several services including customizing results displays. The display illustrates that user pruitt is logged in to MyNCBI. (B) Results are categorized into Tabs. The query for ‘adenylosuccinate lyase’ returns a total of 1545 records (first tab), 715 of which are RefSeq records (last tab). The display illustrates that additional tabs were added to the display to report result subsets for Bacteria and for proteins that have links to the NCBI Map Viewer. (C) Numerous links are calculated between records and can be accessed via the default ‘Links’ menu, or as shown here, the complete set of links can be shown for each record by selecting the option to display links as ‘Plain Links’ in MyNCBI. The link to ‘PubMed (RefSeq)’ returns all publications that are associated with the Entrez Gene record and thus may include a more comprehensive bibliography than that annotated on the RefSeq record.

See this image and copyright information in PMC

References

1. Schuler G.D., Epstein J.A., Ohkawa H., Kans J.A. Entrez: molecular biology database and retrieval system. Methods Enzymol. 1996;266:141–162. - PubMed
1. Altschul S.F., Gish W., Miller W., Myers E.W., Lipman D.J. Basic local alignment search tool. J. Mol. Biol. 1990;215:403–410. - PubMed
1. Altschul S.F., Madden T.L., Schaffer A.A., Zhang J., Zhang Z., Miller W., Lipman D.J. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 1997;25:3389–3402. - PMC - PubMed
1. Benson D.A., Karsch-Mizrachi I., Lipman D.J., Ostell J., Wheeler D.L. GenBank. Nucleic Acids Res. 2007 in press. - PMC - PubMed
1. Wheeler D.L., Barrett T., Benson D.A., Bryant S.H., Canese K., Chetvernin V., Church D.M., DiCuccio M., Edgar R., Federhen S., et al. Database resources of the National Center for Biotechnology Information. Nucleic Acids Res. 2007 in press. - PMC - PubMed

Publication types

Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

Substances

Actions

Grants and funding

Intramural NIH HHS/United States

LinkOut - more resources

Full Text Sources
Other Literature Sources
- The Lens - Patent Citations Database

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Affiliation

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Authors

Affiliation

Abstract

Figures

References

Publication types

MeSH terms

Substances

Grants and funding

LinkOut - more resources

Full Text Sources

Other Literature Sources