Protein structural domains: analysis of the 3Dee domains database
- PMID: 11151005
Protein structural domains: analysis of the 3Dee domains database
Abstract
The 3Dee database of domain definitions was developed as a comprehensive collection of domain definitions for all three-dimensional structures in the Protein Data Bank (PDB). The database includes definitions for complex, multiple-segment and multiple-chain domains as well as simple sequential domains, organized in a structural hierarchy. Two different snapshots of the 3Dee database were analyzed at September 1996 and November 1999. For the November 1999 release, 7,995 PDB entries contained 13,767 protein chains and gave rise to 18,896 domains. The domain sequences clustered into 1,715 domain sequence families, which were further clustered into a conservative 1,199 domain structure families (families with similar folds). The proportion of different domain structure families per domain sequence family increases from 84% for domains 1-100 residues long to 100% for domains greater than 600 residues. This is in keeping with the idea that longer chains will have more alternative folds available to them. Of the representative domains from the domain sequence families, 49% are in the range of 51-150 residues, whereas 64% of the representative chains over 200 residues have more than 1 domain. Of the representative chains, 8.5% are part of multichain domains. The largest multichain domain in the database has 14 chains and 1,400 residues, whereas the largest single-chain domain has 907 residues. The largest number of domains found in a protein is 13. The analysis shows that over the history of the PDB, new domain folds have been discovered at a slower rate than by random selection of all known folds. Between 1992 and 1997, a constant 1 in 11 new domains deposited in the PDB has shown no sequence similarity to a previously known domain sequence family, and only 1 in 15 new domain structures has had a fold that has not been seen previously. A comparison of the September 1996 release of 3Dee to the Structural Classification of Proteins (SCOP) showed that the domain definitions agreed for 80% of the representative protein chains. However, 3Dee provided explicit domain boundaries for more proteins. 3Dee is accessible on the World Wide Web at http://barton.ebi.ac.uk/servers/3Dee.html.
Similar articles
-
3Dee: a database of protein structural domains.Bioinformatics. 2001 Feb;17(2):200-1. doi: 10.1093/bioinformatics/17.2.200. Bioinformatics. 2001. PMID: 11238081
-
Progress of structural genomics initiatives: an analysis of solved target structures.J Mol Biol. 2005 May 20;348(5):1235-60. doi: 10.1016/j.jmb.2005.03.037. Epub 2005 Apr 2. J Mol Biol. 2005. PMID: 15854658
-
Accurate domain identification with structure-anchored hidden Markov models, saHMMs.Proteins. 2009 Aug 1;76(2):343-52. doi: 10.1002/prot.22349. Proteins. 2009. PMID: 19173309
-
Contemporary approaches to protein structure classification.Bioessays. 1998 Nov;20(11):884-91. doi: 10.1002/(SICI)1521-1878(199811)20:11<884::AID-BIES3>3.0.CO;2-H. Bioessays. 1998. PMID: 9872054 Review.
-
Protein folds, functions and evolution.J Mol Biol. 1999 Oct 22;293(2):333-42. doi: 10.1006/jmbi.1999.3054. J Mol Biol. 1999. PMID: 10529349 Review.
Cited by
-
A method for finding candidate conformations for molecular replacement using relative rotation between domains of a known structure.Acta Crystallogr D Biol Crystallogr. 2006 Apr;62(Pt 4):398-409. doi: 10.1107/S0907444906002204. Epub 2006 Mar 18. Acta Crystallogr D Biol Crystallogr. 2006. PMID: 16552141 Free PMC article.
-
Universality in protein residue networks.Biophys J. 2010 Mar 3;98(5):890-900. doi: 10.1016/j.bpj.2009.11.017. Biophys J. 2010. PMID: 20197043 Free PMC article.
-
Estimating the accuracy of protein structures using residual dipolar couplings.J Biomol NMR. 2005 Oct;33(2):83-93. doi: 10.1007/s10858-005-2601-7. J Biomol NMR. 2005. PMID: 16258827
-
A consensus view of fold space: combining SCOP, CATH, and the Dali Domain Dictionary.Protein Sci. 2003 Oct;12(10):2150-60. doi: 10.1110/ps.0306803. Protein Sci. 2003. PMID: 14500873 Free PMC article.
-
The CATH hierarchy revisited-structural divergence in domain superfamilies and the continuity of fold space.Structure. 2009 Aug 12;17(8):1051-62. doi: 10.1016/j.str.2009.06.015. Structure. 2009. PMID: 19679085 Free PMC article.