A Word Pair Dataset for Semantic Similarity and Relatedness in Korean Medical Vocabulary: Reference Development and Validation
- PMID: 34185005
- PMCID: PMC8277378
- DOI: 10.2196/29667
A Word Pair Dataset for Semantic Similarity and Relatedness in Korean Medical Vocabulary: Reference Development and Validation
Abstract
Background: The fact that medical terms require special expertise and are becoming increasingly complex makes it difficult to employ natural language processing techniques in medical informatics. Several human-validated reference standards for medical terms have been developed to evaluate word embedding models using the semantic similarity and relatedness of medical word pairs. However, there are very few reference standards in non-English languages. In addition, because the existing reference standards were developed a long time ago, there is a need to develop an updated standard to represent recent findings in medical sciences.
Objective: We propose a new Korean word pair reference set to verify embedding models.
Methods: From January 2010 to December 2020, 518 medical textbooks, 72,844 health information news, and 15,698 medical research articles were collected, and the top 10,000 medical terms were selected to develop medical word pairs. Attending physicians (n=16) participated in the verification of the developed set with 607 word pairs.
Results: The proportion of word pairs answered by all participants was 90.8% (551/607) for the similarity task and 86.5% (525/605) for the relatedness task. The similarity and relatedness of the word pair showed a high correlation (ρ=0.70, P<.001). The intraclass correlation coefficients to assess the interrater agreements of the word pair sets were 0.47 on the similarity task and 0.53 on the relatedness task. The final reference standard was 604 word pairs for the similarity task and 599 word pairs for relatedness, excluding word pairs with answers corresponding to outliers and word pairs that were answered by less than 50% of all the respondents. When FastText models were applied to the final reference standard word pair sets, the embedding models learning medical documents had a higher correlation between the calculated cosine similarity scores compared to human-judged similarity and relatedness scores (namu, ρ=0.12 vs with medical text for the similarity task, ρ=0.47; namu, ρ=0.02 vs with medical text for the relatedness task, ρ=0.30).
Conclusions: Korean medical word pair reference standard sets for semantic similarity and relatedness were developed based on medical documents from the past 10 years. It is expected that our word pair reference sets will be actively utilized in the development of medical and multilingual natural language processing technology in the future.
Keywords: Korean; fastText; medical word pair; relatedness; similarity; word embedding.
©Yunjin Yum, Jeong Moon Lee, Moon Joung Jang, Yoojoong Kim, Jong-Ho Kim, Seongtae Kim, Unsub Shin, Sanghoun Song, Hyung Joon Joo. Originally published in JMIR Medical Informatics (https://medinform.jmir.org), 24.06.2021.
Conflict of interest statement
Conflicts of Interest: None declared.
Figures





Similar articles
-
Use of word and graph embedding to measure semantic relatedness between Unified Medical Language System concepts.J Am Med Inform Assoc. 2020 Oct 1;27(10):1538-1546. doi: 10.1093/jamia/ocaa136. J Am Med Inform Assoc. 2020. PMID: 33029614 Free PMC article.
-
Corpus domain effects on distributional semantic modeling of medical terms.Bioinformatics. 2016 Dec 1;32(23):3635-3644. doi: 10.1093/bioinformatics/btw529. Epub 2016 Aug 16. Bioinformatics. 2016. PMID: 27531100 Free PMC article.
-
Semantic Similarity and Relatedness between Clinical Terms: An Experimental Study.AMIA Annu Symp Proc. 2010 Nov 13;2010:572-6. AMIA Annu Symp Proc. 2010. PMID: 21347043 Free PMC article.
-
Semantic relatedness and similarity of biomedical terms: examining the effects of recency, size, and section of biomedical publications on the performance of word2vec.BMC Med Inform Decis Mak. 2017 Jul 3;17(1):95. doi: 10.1186/s12911-017-0498-1. BMC Med Inform Decis Mak. 2017. PMID: 28673289 Free PMC article.
-
Predicting Semantic Similarity Between Clinical Sentence Pairs Using Transformer Models: Evaluation and Representational Analysis.JMIR Med Inform. 2021 May 26;9(5):e23099. doi: 10.2196/23099. JMIR Med Inform. 2021. PMID: 34037527 Free PMC article.
Cited by
-
Fine-tuned Sentiment Analysis of COVID-19 Vaccine-Related Social Media Data: Comparative Study.J Med Internet Res. 2022 Oct 17;24(10):e40408. doi: 10.2196/40408. J Med Internet Res. 2022. PMID: 36174192 Free PMC article.
-
A pre-trained BERT for Korean medical natural language processing.Sci Rep. 2022 Aug 16;12(1):13847. doi: 10.1038/s41598-022-17806-8. Sci Rep. 2022. PMID: 35974113 Free PMC article.
References
-
- Safi Z, Abd-Alrazaq A, Khalifa M, Househ M. Technical Aspects of Developing Chatbots for Medical Applications: Scoping Review. J Med Internet Res. 2020 Dec 18;22(12):e19127. doi: 10.2196/19127. https://www.jmir.org/2020/12/e19127/ - DOI - PMC - PubMed
-
- Chen Q, Lee K, Yan S, Kim S, Wei C, Lu Z. BioConceptVec: Creating and evaluating literature-based biomedical concept embeddings on a large scale. PLoS Comput Biol. 2020 Apr;16(4):e1007617. doi: 10.1371/journal.pcbi.1007617. https://dx.plos.org/10.1371/journal.pcbi.1007617 - DOI - DOI - PMC - PubMed
-
- Lee J, Yoon W, Kim S, Kim D, Kim S, So CH, Kang J. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020 Feb 15;36(4):1234–1240. doi: 10.1093/bioinformatics/btz682. http://europepmc.org/abstract/MED/31501885 - DOI - PMC - PubMed
LinkOut - more resources
Full Text Sources
Miscellaneous