Development of a classification scheme for disease-related enzyme information

Carola Söhngen¹, Antje Chang, Dietmar Schomburg

Affiliations

PMID: 21827651
PMCID: PMC3166944
DOI: 10.1186/1471-2105-12-329

Development of a classification scheme for disease-related enzyme information

Carola Söhngen et al. BMC Bioinformatics. 2011.

. 2011 Aug 9:12:329.

doi: 10.1186/1471-2105-12-329.

Authors

Carola Söhngen¹, Antje Chang, Dietmar Schomburg

Affiliation

¹ Technische Universität Braunschweig, Department of Bioinformatics and Biochemistry Langer Kamp 19 B, 38106 Braunschweig, Germany.

PMID: 21827651
PMCID: PMC3166944
DOI: 10.1186/1471-2105-12-329

Abstract

Background: BRENDA (BRaunschweig ENzyme DAtabase, http://www.brenda-enzymes.org) is a major resource for enzyme related information. First and foremost, it provides data which are manually curated from the primary literature. DRENDA (Disease RElated ENzyme information DAtabase) complements BRENDA with a focus on the automatic search and categorization of enzyme and disease related information from title and abstracts of primary publications. In a two-step procedure DRENDA makes use of text mining and machine learning methods.

Results: Currently enzyme and disease related references are biannually updated as part of the standard BRENDA update. 910,897 relations of EC-numbers and diseases were extracted from titles or abstracts and are included in the second release in 2010. The enzyme and disease entity recognition has been successfully enhanced by a further relation classification via machine learning. The classification step has been evaluated by a 5-fold cross validation and achieves an F1 score between 0.802 ± 0.032 and 0.738 ± 0.033 depending on the categories and pre-processing procedures. In the eventual DRENDA content every category reaches a classification specificity of at least 96.7% and a precision that ranges from 86-98% in the highest confidence level, and 64-83% for the smallest confidence level associated with higher recall.

Conclusions: The DRENDA processing chain analyses PubMed, locates references with disease-related information on enzymes and categorises their focus according to the categories causal interaction, therapeutic application, diagnostic usage and ongoing research. The categorisation gives an impression on the focus of the located references. Thus, the relation categorisation can facilitate orientation within the rapidly growing number of references with impact on diseases and enzymes. The DRENDA information is available as additional information in BRENDA.

PubMed Disclaimer

Figures

**Figure 1**
**A schematic illustration of the DRENDA work flow**. The BRENDA enzyme names and synonyms and the MeSH disease terms are used as dictionaries. The PubMed abstracts and titles are searched for co-occurring disease and enzyme entities. A test/train corpus was created for training an SVM and classifying the co-occurrence results according to the categories *causal interaction*, *therapeutic application*, *diagnostic usage* and *ongoing research*. The resulting entries are stored in the DRENDA database.

**Figure 2**
**Receiver operating characteristic (ROC) plots of the models, which achieved the maximal F1 scores**. The ROC plots shown belong to the models, which achieved the maximal F1 scores (table 2) in the five-fold cross-validation with either a *removal* (a) or *replacement* (b) preprocessing applied before the calculation of term weights. The ROC curves are vertical averaged (fixed false positive rates and averages of the corresponding true positive rates of each turn of the five-fold cross validation). In spite of decreasing standard deviation for larger numbers of available training sentences, the largest area under the curve (AUC) is achieved by classifiers for the category *therapeutic application*, which has least annotated sentences in the test/training corpus. See table 2 for the corresponding scalar AUC values of each plot.

**Figure 3**
**The quota of intersection of classification categories (numbers × 10³, rounded)**. The overall amount of distinct EC, disease and PubMed reference combinations in the categories *causal interaction* (grey), *therapeutic application* (blue), *ongoing research* (pink) and *diagnostic usage* (green) in every DRENDA confidence level 1-4. The number of unassigned combinations is listed in the sets (yellow) at the bottom of each plot.

**Figure 4**
**Screen shot of the BRENDA web portal entry with a view on the DRENDA query form**.

**Figure 5**
**Access to the DRENDA data**. The query form (a) provides several fields for entering search pattern information. The fields can be combined arbitrarily for a refinement of the query and meet individual requirements. As an example, a part of the query result table (b) for "Diabetes mellitus" as requested disease and all entries assigned to the category *therapeutic application* with a DRENDA confidence level of 3 and 4.

See this image and copyright information in PMC

References

1. Krallinger M, Leitner F, Valencia A. Analysis of biological processes and diseases using text mining approaches. Methods Mol Biol. 2010;593:341–382. doi: 10.1007/978-1-60327-194-3_16. - DOI - PubMed
1. Jensen LJ, Saric J, Bork P. Literature mining for the biologist: from information retrieval to biological discovery. Nat Rev Genet. 2006;7:119–129. doi: 10.1038/nrg1768. - DOI - PubMed
1. Krallinger M, Izarzugaza JMG, Rodriguez-Penagos C, Valencia A. Extraction of human kinase mutations from literature, databases and genotyping studies. BMC Bioinformatics. 2009;10(Suppl 8):S1. doi: 10.1186/1471-2105-10-S8-S1. - DOI - PMC - PubMed
1. Yeniterzi S, Sezerman U. EnzyMiner: automatic identification of protein level mutations and their impact on target enzymes from PubMed abstracts. BMC Bioinformatics. 2009;10(Suppl 8):S2.. - PMC - PubMed
1. Hoffmann R, Valencia A. A gene network for navigating the literature. Nat Genet. 2004;36:664. doi: 10.1038/ng0704-664. - DOI - PubMed

Publication types

Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
Medical
- MedlinePlus Health Information

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Development of a classification scheme for disease-related enzyme information

Affiliation

Development of a classification scheme for disease-related enzyme information

Authors

Affiliation

Abstract

Figures

References

Publication types

MeSH terms

LinkOut - more resources

Full Text Sources

Medical