Supra-domains: evolutionary units larger than single protein domains
- PMID: 15095989
- DOI: 10.1016/j.jmb.2003.12.026
Supra-domains: evolutionary units larger than single protein domains
Abstract
Domains are the evolutionary units that comprise proteins, and most proteins are built from more than one domain. Domains can be shuffled by recombination to create proteins with new arrangements of domains. Using structural domain assignments, we examined the combinations of domains in the proteins of 131 completely sequenced organisms. We found two-domain and three-domain combinations that recur in different protein contexts with different partner domains. The domains within these combinations have a particular functional and spatial relationship. These units are larger than individual domains and we term them "supra-domains". Amongst the supra-domains, we identified some 1400 (1203 two-domain and 166 three-domain) combinations that are statistically significantly over-represented relative to the occurrence and versatility of the individual component domains. Over one-third of all structurally assigned multi-domain proteins contain these over-represented supra-domains. This means that investigation of the structural and functional relationships of the domains forming these popular combinations would be particularly useful for an understanding of multi-domain protein function and evolution as well as for genome annotation. These and other supra-domains were analysed for their versatility, duplication, their distribution across the three kingdoms of life and their functional classes. By examining the three-dimensional structures of several examples of supra-domains in different biological processes, we identify two basic types of spatial relationships between the component domains: the combined function of the two domains is such that either the geometry of the two domains is crucial and there is a tight constraint on the interface, or the precise orientation of the domains is less important and they are spatially separate. Frequently, the role of the supra-domain becomes clear only once the three-dimensional structure is known. Since this is the case for only a quarter of the supra-domains, we provide a list of the most important unknown supra-domains as potential targets for structural genomics projects.
Similar articles
-
The relationship between domain duplication and recombination.J Mol Biol. 2005 Feb 11;346(1):355-65. doi: 10.1016/j.jmb.2004.11.050. Epub 2004 Dec 23. J Mol Biol. 2005. PMID: 15663950
-
The geometry of domain combination in proteins.J Mol Biol. 2002 Jan 25;315(4):927-39. doi: 10.1006/jmbi.2001.5288. J Mol Biol. 2002. PMID: 11812158
-
Domain combinations in archaeal, eubacterial and eukaryotic proteomes.J Mol Biol. 2001 Jul 6;310(2):311-25. doi: 10.1006/jmbi.2001.4776. J Mol Biol. 2001. PMID: 11428892
-
The many faces of the helix-turn-helix domain: transcription regulation and beyond.FEMS Microbiol Rev. 2005 Apr;29(2):231-62. doi: 10.1016/j.femsre.2004.12.008. FEMS Microbiol Rev. 2005. PMID: 15808743 Review.
-
Structure, function and evolution of multidomain proteins.Curr Opin Struct Biol. 2004 Apr;14(2):208-16. doi: 10.1016/j.sbi.2004.03.011. Curr Opin Struct Biol. 2004. PMID: 15093836 Review.
Cited by
-
Protein domain recurrence and order can enhance prediction of protein functions.Bioinformatics. 2012 Sep 15;28(18):i444-i450. doi: 10.1093/bioinformatics/bts398. Bioinformatics. 2012. PMID: 22962465 Free PMC article.
-
Identification of potential drug targets implicated in Parkinson's disease from human genome: insights of using fused domains in hypothetical proteins as probes.ISRN Neurol. 2011;2011:265253. doi: 10.5402/2011/265253. Epub 2011 Sep 4. ISRN Neurol. 2011. PMID: 22389811 Free PMC article.
-
Improvement in Protein Domain Identification Is Reached by Breaking Consensus, with the Agreement of Many Profiles and Domain Co-occurrence.PLoS Comput Biol. 2016 Jul 29;12(7):e1005038. doi: 10.1371/journal.pcbi.1005038. eCollection 2016 Jul. PLoS Comput Biol. 2016. PMID: 27472895 Free PMC article.
-
Genomic scale sub-family assignment of protein domains.Nucleic Acids Res. 2006 Jul 28;34(13):3625-33. doi: 10.1093/nar/gkl484. Print 2006. Nucleic Acids Res. 2006. PMID: 16877569 Free PMC article.
-
Comparative analysis of barophily-related amino acid content in protein domains of Pyrococcus abyssi and Pyrococcus furiosus.Archaea. 2013;2013:680436. doi: 10.1155/2013/680436. Epub 2013 Sep 26. Archaea. 2013. PMID: 24187517 Free PMC article.
Publication types
MeSH terms
Substances
LinkOut - more resources
Full Text Sources
Other Literature Sources