. 2009 Apr 13:10:107.

doi: 10.1186/1471-2105-10-107.

EFICAz2: enzyme function inference by a combined approach enhanced by machine learning

Adrian K Arakaki¹, Ying Huang, Jeffrey Skolnick

Affiliations

PMID: 19361344
PMCID: PMC2670841
DOI: 10.1186/1471-2105-10-107

EFICAz2: enzyme function inference by a combined approach enhanced by machine learning

Adrian K Arakaki et al. BMC Bioinformatics. 2009.

. 2009 Apr 13:10:107.

doi: 10.1186/1471-2105-10-107.

Authors

Adrian K Arakaki¹, Ying Huang, Jeffrey Skolnick

Affiliation

¹ Center for the Study of Systems Biology, School of Biology, Georgia Institute of Technology, Atlanta, Georgia 30318, USA. adrian.arakaki@gatech.edu

PMID: 19361344
PMCID: PMC2670841
DOI: 10.1186/1471-2105-10-107

Abstract

Background: We previously developed EFICAz, an enzyme function inference approach that combines predictions from non-completely overlapping component methods. Two of the four components in the original EFICAz are based on the detection of functionally discriminating residues (FDRs). FDRs distinguish between member of an enzyme family that are homofunctional (classified under the EC number of interest) or heterofunctional (annotated with another EC number or lacking enzymatic activity). Each of the two FDR-based components is associated to one of two specific kinds of enzyme families. EFICAz exhibits high precision performance, except when the maximal test to training sequence identity (MTTSI) is lower than 30%. To improve EFICAz's performance in this regime, we: i) increased the number of predictive components and ii) took advantage of consensual information from the different components to make the final EC number assignment.

Results: We have developed two new EFICAz components, analogs to the two FDR-based components, where the discrimination between homo and heterofunctional members is based on the evaluation, via Support Vector Machine models, of all the aligned positions between the query sequence and the multiple sequence alignments associated to the enzyme families. Benchmark results indicate that: i) the new SVM-based components outperform their FDR-based counterparts, and ii) both SVM-based and FDR-based components generate unique predictions. We developed classification tree models to optimally combine the results from the six EFICAz components into a final EC number prediction. The new implementation of our approach, EFICAz2, exhibits a highly improved prediction precision at MTTSI < 30% compared to the original EFICAz, with only a slight decrease in prediction recall. A comparative analysis of enzyme function annotation of the human proteome by EFICAz2 and KEGG shows that: i) when both sources make EC number assignments for the same protein sequence, the assignments tend to be consistent and ii) EFICAz2 generates considerably more unique assignments than KEGG.

Conclusion: Performance benchmarks and the comparison with KEGG demonstrate that EFICAz2 is a powerful and precise tool for enzyme function annotation, with multiple applications in genome analysis and metabolic pathway reconstruction. The EFICAz2 web service is available at: http://cssb.biology.gatech.edu/skolnick/webservice/EFICAz2/index.html.

PubMed Disclaimer

Figures

**Figure 1**
**Prediction performance of the FDR-based and SVM-based approaches applied to Multiple Pfam enzyme families**. For three-field (A, B) or four-field EC number classifiers (C, D), the average recall (A, C) and average precision (B, D) of the FDR-based (blue columns) and SVM-based (red columns) approaches is plotted at different intervals of maximal test to training sequence identity (MTTSI). The average of each performance indicator is done over all the EC numbers defined in the specified MTTSI interval (numbers at the bottom of each column). Details about the benchmark can be found in "Benchmarking of EFICAz²version 10", in the Methods section. Statistically significant differences in performance are indicated by black lines under the corresponding columns (see "Statistical analyses", in the Methods section). Values on top of each column represent average +/- standard deviation.

**Figure 2**
**Prediction performance of the FDR-based and SVM-based approaches applied to CHIEFc enzyme families**. For three-field (A, B) or four-field EC number classifiers (C, D), the average recall (A, C) and average precision (B, D) of the FDR-based (blue columns) and SVM-based (red columns) approaches is plotted at different intervals of maximal test to training sequence identity (MTTSI). The average of each performance indicator is done over all the EC numbers defined in the specified MTTSI interval (numbers at the bottom of each column). Details about the benchmark can be found in "Benchmarking of EFICAz²version 10", in the Methods section. Statistically significant differences in performance are indicated by black lines under the corresponding columns (see "Statistical analyses", in the Methods section). Values on top of each column represent average +/- standard deviation.

**Figure 3**
**Prediction overlap of FDR-based and SVM-based methods**. The fractions of test sequences (corresponding to the benchmark described in "Benchmarking of EFICAz²version 10", in the Methods section) correctly predicted by three or four-field EC number classifiers applied to Multiple Pfam or CHIEFc enzyme families are represented. For combination of enzyme family and level of description of the classifiers, we show the fraction corresponding to unique predictions made by the FDR-based (blue) or SVM-based method (green), and the fraction corresponding to predictions made by both (orange) or none of the methods (yellow).

**Figure 4**
**Prediction performance of different EFICAz implementations**. For three-field (A, B) or four-field EC number classifiers (C, D), the average recall (A, C) and average precision (B, D) of the original EFICAz (green columns), EFICAz plus the new SVM-based components (blue columns) and EFICAz²(red columns) is plotted at different intervals of maximal test to training sequence identity (MTTSI). The average of each performance indicator is done over all the EC numbers defined in the specified MTTSI interval (numbers at the bottom of each column). Details about the benchmark can be found in "Benchmarking of EFICAz²version 10", in the Methods section. Statistically significant differences in performance are indicated by black lines under the corresponding columns (see "Statistical analyses", in the Methods section). Values on top of each column represent average +/- standard deviation.

**Figure 5**
**Predictive models for EFICAz²based on classification trees**. Classification trees corresponding to three-field (A, B) and four-field EC numbers (C, D) to integrate predictions from each of the six EFICAz²components for protein sequences that exhibit MTTSI < 30% (A, C) or MTTSI ≥ 30% (B, D). CH_FDR= CHIEFc family based FDR recognition; PF_FDR= Multiple Pfam family based FDR recognition; CH_SIT= CHIEFc family specific SIT evaluation; Prst = High specificity multiple PROSITE pattern recognition; CH_svm= CHIEFc family based SVM evaluation; PF_svm= Multiple Pfam family based SVM evaluation.

**Figure 6**
**Distribution of the number of test sequences per enzyme type**. Distribution of 9,397 test enzyme sequences into 145 types of three-field EC numbers (green columns) and 6,996 test enzyme sequences into 614 types of four-field EC numbers (red columns).

See this image and copyright information in PMC

Cited by

Horizontal gene transfer and genome evolution in Methanosarcina.
Garushyants SK, Kazanov MD, Gelfand MS. Garushyants SK, et al. BMC Evol Biol. 2015 Jun 5;15:102. doi: 10.1186/s12862-015-0393-2. BMC Evol Biol. 2015. PMID: 26044078 Free PMC article.
The use of evolutionary patterns in protein annotation.
Wilkins AD, Bachman BJ, Erdin S, Lichtarge O. Wilkins AD, et al. Curr Opin Struct Biol. 2012 Jun;22(3):316-25. doi: 10.1016/j.sbi.2012.05.001. Epub 2012 May 24. Curr Opin Struct Biol. 2012. PMID: 22633559 Free PMC article. Review.
Reference genomes and transcriptomes of Nicotiana sylvestris and Nicotiana tomentosiformis.
Sierro N, Battey JN, Ouadi S, Bovet L, Goepfert S, Bakaher N, Peitsch MC, Ivanov NV. Sierro N, et al. Genome Biol. 2013 Jun 17;14(6):R60. doi: 10.1186/gb-2013-14-6-r60. Genome Biol. 2013. PMID: 23773524 Free PMC article.
Comparative genomics of cell envelope components in mycobacteria.
Banerjee R, Vats P, Dahale S, Kasibhatla SM, Joshi R. Banerjee R, et al. PLoS One. 2011 May 6;6(5):e19280. doi: 10.1371/journal.pone.0019280. PLoS One. 2011. PMID: 21573108 Free PMC article.
Genomics of Loa loa, a Wolbachia-free filarial parasite of humans.
Desjardins CA, Cerqueira GC, Goldberg JM, Dunning Hotopp JC, Haas BJ, Zucker J, Ribeiro JM, Saif S, Levin JZ, Fan L, Zeng Q, Russ C, Wortman JR, Fink DL, Birren BW, Nutman TB. Desjardins CA, et al. Nat Genet. 2013 May;45(5):495-500. doi: 10.1038/ng.2585. Epub 2013 Mar 24. Nat Genet. 2013. PMID: 23525074 Free PMC article.

See all "Cited by" articles

References

1. Arakaki AK, Tian W, Skolnick J. High precision multi-genome scale reannotation of enzyme function by EFICAz. BMC Genomics. 2006;7:315. doi: 10.1186/1471-2164-7-315. - DOI - PMC - PubMed
1. Freilich S, Spriggs RV, George RA, Al-Lazikani B, Swindells M, Thornton JM. The complement of enzymatic sets in different species. J Mol Biol. 2005;349:745–763. doi: 10.1016/j.jmb.2005.04.027. - DOI - PubMed
1. Webb EC. Enzyme nomenclature 1992: recommendations of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology on the nomenclature and classification of enzymes. San Diego: Published for the International Union of Biochemistry and Molecular Biology by Academic Press; 1992.
1. Glasner ME, Gerlt JA, Babbitt PC. Evolution of enzyme superfamilies. Curr Opin Chem Biol. 2006;10:492–497. doi: 10.1016/j.cbpa.2006.08.012. - DOI - PubMed
1. Ginsburg H. Caveat emptor: limitations of the automated reconstruction of metabolic pathways in Plasmodium. Trends Parasitol. 2008;25:37–43. doi: 10.1016/j.pt.2008.08.012. - DOI - PubMed

Publication types

Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

Substances

Actions

Grants and funding

LinkOut - more resources

Full Text Sources
Molecular Biology Databases
- NIAID Data Ecosystem - Find datasets on Infectious and Immune-mediated Diseases

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

EFICAz2: enzyme function inference by a combined approach enhanced by machine learning

Affiliation

EFICAz2: enzyme function inference by a combined approach enhanced by machine learning

Authors

Affiliation

Abstract

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

Substances

Grants and funding

LinkOut - more resources

Full Text Sources

Molecular Biology Databases