The contribution of co-reference resolution to supervised relation detection between bacteria and biotopes entities
- PMID: 26201352
- PMCID: PMC4511182
- DOI: 10.1186/1471-2105-16-S10-S6
The contribution of co-reference resolution to supervised relation detection between bacteria and biotopes entities
Abstract
Background: The acquisition of knowledge about relations between bacteria and their locations (habitats and geographical locations) in short texts about bacteria, as defined in the BioNLP-ST 2013 Bacteria Biotope task, depends on the detection of co-reference links between mentions of entities of each of these three types. To our knowledge, no participant in this task has investigated this aspect of the situation. The present work specifically addresses issues raised by this situation: (i) how to detect these co-reference links and associated co-reference chains; (ii) how to use them to prepare positive and negative examples to train a supervised system for the detection of relations between entity mentions; (iii) what context around which entity mentions contributes to relation detection when co-reference chains are provided.
Results: We present experiments and results obtained both with gold entity mentions (task 2 of BioNLP-ST 2013) and with automatically detected entity mentions (end-to-end system, in task 3 of BioNLP-ST 2013). Our supervised mention detection system uses a linear chain Conditional Random Fields classifier, and our relation detection system relies on a Logistic Regression (aka Maximum Entropy) classifier. They use a set of morphological, morphosyntactic and semantic features. To minimize false inferences, co-reference resolution applies a set of heuristic rules designed to optimize precision. They take into account the types of the detected entity mentions, and take advantage of the didactic nature of the texts of the corpus, where a large proportion of bacteria naming is fairly explicit (although natural referring expressions such as "the bacteria" are common). The resulting system achieved a 0.495 F-measure on the official test set when taking as input the gold entity mentions, and a 0.351 F-measure when taking as input entity mentions predicted by our CRF system, both of which are above the best BioNLP-ST 2013 participant system.
Conclusions: We show that co-reference resolution substantially improves over a baseline system which does not use co-reference information: about 3.5 F-measure points on the test corpus for the end-to-end system (5.5 points on the development corpus) and 7 F-measure points on both development and test corpora when gold mentions are used. While this outperforms the best published system on the BioNLP-ST 2013 Bacteria Biotope dataset, we consider that it provides mostly a stronger baseline from which more work can be started. We also emphasize the importance and difficulty of designing a comprehensive gold standard co-reference annotation, which we explain is a key point to further progress on the task.
Figures







Similar articles
-
Overview of the gene regulation network and the bacteria biotope tasks in BioNLP'13 shared task.BMC Bioinformatics. 2015;16 Suppl 10(Suppl 10):S1. doi: 10.1186/1471-2105-16-S10-S1. Epub 2015 Jul 13. BMC Bioinformatics. 2015. PMID: 26202448 Free PMC article.
-
Detection and categorization of bacteria habitats using shallow linguistic analysis.BMC Bioinformatics. 2015;16 Suppl 10(Suppl 10):S5. doi: 10.1186/1471-2105-16-S10-S5. Epub 2015 Jul 13. BMC Bioinformatics. 2015. PMID: 26201262 Free PMC article.
-
Linking entities through an ontology using word embeddings and syntactic re-ranking.BMC Bioinformatics. 2019 Mar 27;20(1):156. doi: 10.1186/s12859-019-2678-8. BMC Bioinformatics. 2019. PMID: 30917789 Free PMC article.
-
Community challenges in biomedical text mining over 10 years: success, failure and the future.Brief Bioinform. 2016 Jan;17(1):132-44. doi: 10.1093/bib/bbv024. Epub 2015 May 1. Brief Bioinform. 2016. PMID: 25935162 Free PMC article. Review.
-
Semantic annotation in biomedicine: the current landscape.J Biomed Semantics. 2017 Sep 22;8(1):44. doi: 10.1186/s13326-017-0153-x. J Biomed Semantics. 2017. PMID: 28938912 Free PMC article. Review.
Cited by
-
Extracting medications and associated adverse drug events using a natural language processing system combining knowledge base and deep learning.J Am Med Inform Assoc. 2020 Jan 1;27(1):56-64. doi: 10.1093/jamia/ocz141. J Am Med Inform Assoc. 2020. PMID: 31591641 Free PMC article.
-
Extraction of Information Related to Adverse Drug Events from Electronic Health Record Notes: Design of an End-to-End Model Based on Deep Learning.JMIR Med Inform. 2018 Nov 26;6(4):e12159. doi: 10.2196/12159. JMIR Med Inform. 2018. PMID: 30478023 Free PMC article.
-
Relation Extraction from Clinical Narratives Using Pre-trained Language Models.AMIA Annu Symp Proc. 2020 Mar 4;2019:1236-1245. eCollection 2019. AMIA Annu Symp Proc. 2020. PMID: 32308921 Free PMC article.
-
Bridging semantics and syntax with graph algorithms-state-of-the-art of extracting biomedical relations.Brief Bioinform. 2017 Jan;18(1):160-178. doi: 10.1093/bib/bbw001. Epub 2016 Feb 5. Brief Bioinform. 2017. PMID: 26851224 Free PMC article.
-
A neural joint model for entity and relation extraction from biomedical text.BMC Bioinformatics. 2017 Mar 31;18(1):198. doi: 10.1186/s12859-017-1609-9. BMC Bioinformatics. 2017. PMID: 28359255 Free PMC article.
References
-
- Kim JD, Pyysalo S, Ohta T, Bossy R, Nguyen N, Tsujii J. BioNLP Shared Task 2011 Workshop Proc. Portland, OR: ACL; 2011. Overview of BioNLP Shared Task 2011; pp. 1–6.
-
- Nédellec C, Bossy R, Kim JD, Kim JJ, Ohta T, Pyysalo S, Zweigenbaum P. Proceedings of the BioNLP Shared Task 2013 Workshop. Sofia, Bulgaria: Association for Computational Linguistics; 2013. Overview of BioNLP Shared Task 2013; pp. 1–7.
-
- Coreference task definition (v2.3) In: Proceedings of the Sixth Message Understanding Conference (MUC-6) Columbia, MD . 1995. pp. 335–344.
Publication types
MeSH terms
LinkOut - more resources
Full Text Sources