Simulation and annotation of global acronyms

Maxim Filimonov¹, Daphné Chopard¹, Irena Spasić¹

Affiliations

PMID: 35482480
PMCID: PMC9154234
DOI: 10.1093/bioinformatics/btac298

Simulation and annotation of global acronyms

Maxim Filimonov et al. Bioinformatics. 2022.

. 2022 May 26;38(11):3136-3138.

doi: 10.1093/bioinformatics/btac298.

Authors

Maxim Filimonov¹, Daphné Chopard¹, Irena Spasić¹

Affiliation

¹ School of Computer Science and Informatics, Cardiff University, Cardiff CF24 4AG, UK.

PMID: 35482480
PMCID: PMC9154234
DOI: 10.1093/bioinformatics/btac298

Abstract

Motivation: Global acronyms are used in written text without their formal definitions. This makes it difficult to automatically interpret their sense as acronyms tend to be ambiguous. Supervised machine learning approaches to sense disambiguation require large training datasets. In clinical applications, large datasets are difficult to obtain due to patient privacy. Manual data annotation creates an additional bottleneck.

Results: We proposed an approach to automatically modifying scientific abstracts to (i) simulate global acronym usage and (ii) annotate their senses without the need for external sources or manual intervention. We implemented it as a web-based application, which can create large datasets that in turn can be used to train supervised approaches to word sense disambiguation of biomedical acronyms.

Availability and implementation: The datasets will be generated on demand based on a user query and will be downloadable from https://datainnovation.cardiff.ac.uk/acronyms/.

PubMed Disclaimer

Figures

See this image and copyright information in PMC

References

1. Agirre E., Stevenson M. (2006) Knowledge sources for WSD. In: Agirre, E. and Edmonds, P. (eds) Word Sense Disambiguation, pp. 217–251. https://doi.org/10.1007/978-1-4020-4809-8_8.
1. Ao H., Takagi T. (2005) ALICE: an algorithm to extract abbreviations from MEDLINE. J. Am. Med. Inform. Assoc., 12, 576–586. - PMC - PubMed
1. Bodenreider O. (2004) The unified medical language system (UMLS): integrating biomedical terminology. Nucleic Acids Res., 32, D267–D270. - PMC - PubMed
1. Chang J.T. et al. (2002) Creating an online dictionary of abbreviations from MEDLINE. J. Am. Med. Inform. Assoc., 9, 612–620. - PMC - PubMed
1. Devlin J. et al. (2018) BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1, Association for Computational Linguistics, Minneapolis, Minnesota, pp. 4171–4186.

LinkOut - more resources

Full Text Sources

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Simulation and annotation of global acronyms

Affiliation

Simulation and annotation of global acronyms

Authors

Affiliation

Abstract

Figures

References

LinkOut - more resources

Full Text Sources