Comparative Study

. 2001 Oct;11(10):1632-40.

doi: 10.1101/gr.183801.

Annotation transfer for genomics: measuring functional divergence in multi-domain proteins

H Hegyi¹, M Gerstein

Affiliations

PMID: 11591640
PMCID: PMC311165
DOI: 10.1101/gr.183801

Comparative Study

Annotation transfer for genomics: measuring functional divergence in multi-domain proteins

H Hegyi et al. Genome Res. 2001 Oct.

. 2001 Oct;11(10):1632-40.

doi: 10.1101/gr.183801.

Authors

H Hegyi¹, M Gerstein

Affiliation

¹ Department of Molecular Biophysics and Biochemistry, Yale University, New Haven, Connecticut 06520, USA.

PMID: 11591640
PMCID: PMC311165
DOI: 10.1101/gr.183801

Abstract

Annotation transfer is a principal process in genome annotation. It involves "transferring" structural and functional annotation to uncharacterized open reading frames (ORFs) in a newly completed genome from experimentally characterized proteins similar in sequence. To prevent errors in genome annotation, it is important that this process be robust and statistically well-characterized, especially with regard to how it depends on the degree of sequence similarity. Previously, we and others have analyzed annotation transfer in single-domain proteins. Multi-domain proteins, which make up the bulk of the ORFs in eukaryotic genomes, present more complex issues in functional conservation. Here we present a large-scale survey of annotation transfer in these proteins, using scop superfamilies to define domain folds and a thesaurus based on SWISS-PROT keywords to define functional categories. Our survey reveals that multi-domain proteins have significantly less functional conservation than single-domain ones, except when they share the exact same combination of domain folds. In particular, we find that for multi-domain proteins, approximate function can be accurately transferred with only 35% certainty for pairs of proteins sharing one structural superfamily. In contrast, this value is 67% for pairs of single-domain proteins sharing the same structural superfamily. On the other hand, if two multi-domain proteins contain the same combination of two structural superfamilies the probability of their sharing the same function increases to 80% in the case of complete coverage along the full length of both proteins, this value increases further to > 90%. Moreover, we found that only 70 of the current total of 455 structural superfamilies are found in both single and multi-domain proteins and only 14 of these were associated with the same function in both categories of proteins. We also investigated the degree to which function could be transferred between pairs of multi-domain proteins with respect to the degree of sequence similarity between them, finding that functional divergence at a given amount of sequence similarity is always about two-fold greater for pairs of multi-domain proteins (sharing similarity over a single domain) in comparison to pairs of single-domain ones, though the overall shape of the relationship is quite similar. Further information is available at http://partslist.org/func or http://bioinfo.mbb.yale.edu/partslist/func.

PubMed Disclaimer

Figures

**Figure 1**
Schematic illustrating annotation transfer. This figure illustrates the process of annotation transfer for a group of hypothetical TIM barrel proteins. The *leftmost* panel represents sequence comparisons between idealized barrel domains from a number of organisms. The next panel shows analogous results for structural comparison, and the panel after that, functional comparison. The *rightmost* panel represents sequence comparisons between idealized multi-domain proteins that match over a single domain, the subject of much of this paper.

**Figure 3**
Distribution of proteins amongst broad structural and functional classes; the distribution of the matches among the seven structural and two functional classes in single- and multi-domain proteins. The single-domain and multi-domain matches each total 100%, independently of each other. The horizontal axis indicates the seven scop classes, which are (from 1 to 7): all-alpha, all-beta, alpha/beta, alpha + beta, multi-domain, membrane, and small protein.

**Figure 4**
Divergence in function with respect to sequence similarity. Relative number of matching domains with multiple functions, as the function of e-value threshold. Diamonds represent single-domain proteins, squares multi-domain ones (matching just for a single domain), respectively. The first value on the X-axis starts at 4 (corresponding to an e-value=10⁻⁴).

See this image and copyright information in PMC

Cited by

Bacterial regulatory networks are extremely flexible in evolution.
Lozada-Chávez I, Janga SC, Collado-Vides J. Lozada-Chávez I, et al. Nucleic Acids Res. 2006 Jul 13;34(12):3434-45. doi: 10.1093/nar/gkl423. Print 2006. Nucleic Acids Res. 2006. PMID: 16840530 Free PMC article.
LigProf: a simple tool for in silico prediction of ligand-binding sites.
Koczyk G, Wyrwicz LS, Rychlewski L. Koczyk G, et al. J Mol Model. 2007 Mar;13(3):445-55. doi: 10.1007/s00894-006-0165-4. Epub 2007 Jan 3. J Mol Model. 2007. PMID: 17200839
Annotation transfer between genomes: protein-protein interologs and protein-DNA regulogs.
Yu H, Luscombe NM, Lu HX, Zhu X, Xia Y, Han JD, Bertin N, Chung S, Vidal M, Gerstein M. Yu H, et al. Genome Res. 2004 Jun;14(6):1107-18. doi: 10.1101/gr.1774904. Genome Res. 2004. PMID: 15173116 Free PMC article.
Diversity in the architecture of ATLs, a family of plant ubiquitin-ligases, leads to recognition and targeting of substrates in different cellular environments.
Aguilar-Hernández V, Aguilar-Henonin L, Guzmán P. Aguilar-Hernández V, et al. PLoS One. 2011;6(8):e23934. doi: 10.1371/journal.pone.0023934. Epub 2011 Aug 24. PLoS One. 2011. PMID: 21887349 Free PMC article.
Analyses of domains and domain fusions in human proto-oncogenes.
Liu Q, Huang J, Liu H, Wan P, Ye X, Xu Y. Liu Q, et al. BMC Bioinformatics. 2009 Mar 17;10:88. doi: 10.1186/1471-2105-10-88. BMC Bioinformatics. 2009. PMID: 19292927 Free PMC article.

See all "Cited by" articles

References

1. Altschul S F, Madden T L, Schaffer A A, Zhang J, Zhang Z, Miller W, Lipman D J. Gapped BLAST and PSI-BLAST: A new generation of protein database search programs. Nucleic Acids Res. 1997;25:3389–3402. - PMC - PubMed
1. Bairoch A. The ENZYME database in 2000. Nucleic Acids Res. 2000;28:304–5. - PMC - PubMed
1. Bairoch A, Apweiler R. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000. Nucleic Acids Res. 2000;28:45–8. - PMC - PubMed
1. Chothia C, Lesk A M. The relation between the divergence of sequence and structure in proteins. EMBO J. 1986;5:823–826. - PMC - PubMed
1. Devos D, Valencia A. Practical limits of function prediction. Proteins. 2000;41:98–107. - PubMed

Publication types

Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Annotation transfer for genomics: measuring functional divergence in multi-domain proteins

Affiliation

Annotation transfer for genomics: measuring functional divergence in multi-domain proteins

Authors

Affiliation

Abstract

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

LinkOut - more resources

Full Text Sources

Abstract

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

Related information

LinkOut - more resources

Full Text Sources