Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 1996 Nov;112(3):1177-83.
doi: 10.1104/pp.112.3.1177.

The construction of Arabidopsis expressed sequence tag assemblies. A new resource to facilitate gene identification

Affiliations

The construction of Arabidopsis expressed sequence tag assemblies. A new resource to facilitate gene identification

S D Rounsley et al. Plant Physiol. 1996 Nov.

Abstract

The generation of large numbers of partial cDNA sequences, or expressed sequence tags (ESTs), has provided a method with which to sample a large number of genes from an organism. More than 25,000 Arabidopsis thaliana ESTs have been deposited in public databases, producing the largest collection of ESTs for any plant species. We describe here the application of a method of reducing redundancy and increasing information content in this collection by grouping overlapping ESTs representing the same gene into a "contig" or assembly. The increased information content of these assemblies allows more putative identifications to be assigned based on the results of similarity searches with nucleotide and protein databases. The results of this analysis indicate that sequence information is available for approximately 12,600 nonoverlapping ESTs from Arabidopsis. Comparison of the assemblies with 953 Arabidopsis coding sequences indicates that up to 57% of all Arabidopsis genes are represented by an EST. Clustering analysis of these sequences suggests that between 300 and 700 gene families are represented by between 700 and 2000 sequences in the EST database. A database of the assembled sequences, their putative identifications, and cellular roles is available through the World Wide Web.

PubMed Disclaimer

References

    1. J Mol Biol. 1981 Mar 25;147(1):195-7 - PubMed
    1. Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 - PubMed
    1. J Mol Biol. 1990 Oct 5;215(3):403-10 - PubMed
    1. Proc Natl Acad Sci U S A. 1991 Mar 1;88(5):1943-7 - PubMed
    1. Science. 1991 Jun 21;252(5013):1651-6 - PubMed

Publication types