Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study

Zoltan P Majdik¹, S Scott Graham², Jade C Shiva Edward², Sabrina N Rodriguez³, Martha S Karnes⁴, Jared T Jensen², Joshua B Barbour⁵, Justin F Rousseau^{6

7}

Affiliations

¹ Department of Communication, North Dakota State University, Fargo, ND, United States.
² Department of Rhetoric & Writing, The University of Texas at Austin, Austin, TX, United States.
³ Department of Neurology, The Dell Medical School, The University of Texas at Austin, Austin, TX, United States.
⁴ Department of Rhetoric & Writing, University of Arkansas Little Rock, Little Rock, AR, United States.
⁵ Department of Communication, The University of Illinois at Urbana-Champaign, Urbana, IL, United States.
⁶ Statistical Planning and Analysis Section, Department of Neurology, The University of Texas Southwestern Medical Center, Dallas, TX, United States.
⁷ Peter O'Donnell Jr. Brain Institute, The University of Texas Southwestern Medical Center, Dallas, TX, United States.

PMID: 38875593
PMCID: PMC11140272
DOI: 10.2196/52095

Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study

Zoltan P Majdik et al. JMIR AI. 2024.

. 2024 May 16:3:e52095.

doi: 10.2196/52095.

Authors

Zoltan P Majdik¹, S Scott Graham², Jade C Shiva Edward², Sabrina N Rodriguez³, Martha S Karnes⁴, Jared T Jensen², Joshua B Barbour⁵, Justin F Rousseau^{6

7}

Affiliations

¹ Department of Communication, North Dakota State University, Fargo, ND, United States.
² Department of Rhetoric & Writing, The University of Texas at Austin, Austin, TX, United States.
³ Department of Neurology, The Dell Medical School, The University of Texas at Austin, Austin, TX, United States.
⁴ Department of Rhetoric & Writing, University of Arkansas Little Rock, Little Rock, AR, United States.
⁵ Department of Communication, The University of Illinois at Urbana-Champaign, Urbana, IL, United States.
⁶ Statistical Planning and Analysis Section, Department of Neurology, The University of Texas Southwestern Medical Center, Dallas, TX, United States.
⁷ Peter O'Donnell Jr. Brain Institute, The University of Texas Southwestern Medical Center, Dallas, TX, United States.

PMID: 38875593
PMCID: PMC11140272
DOI: 10.2196/52095

Abstract

Background: Large language models (LLMs) have the potential to support promising new applications in health informatics. However, practical data on sample size considerations for fine-tuning LLMs to perform specific tasks in biomedical and health policy contexts are lacking.

Objective: This study aims to evaluate sample size and sample selection techniques for fine-tuning LLMs to support improved named entity recognition (NER) for a custom data set of conflicts of interest disclosure statements.

Methods: A random sample of 200 disclosure statements was prepared for annotation. All "PERSON" and "ORG" entities were identified by each of the 2 raters, and once appropriate agreement was established, the annotators independently annotated an additional 290 disclosure statements. From the 490 annotated documents, 2500 stratified random samples in different size ranges were drawn. The 2500 training set subsamples were used to fine-tune a selection of language models across 2 model architectures (Bidirectional Encoder Representations from Transformers [BERT] and Generative Pre-trained Transformer [GPT]) for improved NER, and multiple regression was used to assess the relationship between sample size (sentences), entity density (entities per sentence [EPS]), and trained model performance (F₁-score). Additionally, single-predictor threshold regression models were used to evaluate the possibility of diminishing marginal returns from increased sample size or entity density.

Results: Fine-tuned models ranged in topline NER performance from F₁-score=0.79 to F₁-score=0.96 across architectures. Two-predictor multiple linear regression models were statistically significant with multiple R² ranging from 0.6057 to 0.7896 (all P<.001). EPS and the number of sentences were significant predictors of F₁-scores in all cases ( P<.001), except for the GPT-2_large model, where EPS was not a significant predictor (P=.184). Model thresholds indicate points of diminishing marginal return from increased training data set sample size measured by the number of sentences, with point estimates ranging from 439 sentences for RoBERTa_large to 527 sentences for GPT-2_large. Likewise, the threshold regression models indicate a diminishing marginal return for EPS with point estimates between 1.36 and 1.38.

Conclusions: Relatively modest sample sizes can be used to fine-tune LLMs for NER tasks applied to biomedical text, and training data entity density should representatively approximate entity density in production data. Training data quality and a model architecture's intended use (text generation vs text processing or classification) may be as, or more, important as training data volume and model parameter size.

Keywords: annotation; conflict of interest; disclosure; disclosures; expert annotation; fine-tuning; language model; large language models; machine learning; named-entity recognition; natural language processing; sample; sample size; statement; statements; transfer learning.

©Zoltan P Majdik, S Scott Graham, Jade C Shiva Edward, Sabrina N Rodriguez, Martha S Karnes, Jared T Jensen, Joshua B Barbour, Justin F Rousseau. Originally published in JMIR AI (https://ai.jmir.org), 16.05.2024.

PubMed Disclaimer

Conflict of interest statement

Conflicts of Interest: SSG reports grant funding from National Institute of General Medical Sciences (NIGMS) and the Texas Health and Human Services Commission. ZPM reports grant funding from NIGMS and National Science Foundation. SNR reports grant funding from the National Institute of Neurological Disorders and Stroke. JBB reports grant funding from NIGMS, National Science Foundation, and Blue Cross Blue Shield/Health Care Service Corporation. JRF reports grant funding from NIGMS, National Institute of Mental Health (NIMH), National Institute of Allergy and Infectious Diseases (NIAID), National Library of Medicine (NLM), Health Care Cost Institute, Austin Public Health, Texas Child Mental Health Care Consortium, Texas Alzheimer Research and Care Consortium, and the Michael & Susan Dell Foundation. JFR also reports receiving a grant from the NIH Division of Loan Repayment. All other authors report no conflicts of interest.

Figures

**Figure 1**
Single predictor plots for the number of sentence (left) and EPS (right). Fit with a generalized additive model. EPS: entities per sentence.

See this image and copyright information in PMC

References

1. Nadkarni PM, Ohno-Machado L, Chapman WW. Natural language processing: an introduction. J Am Med Inform Assoc. 2011;18(5):544–551. doi: 10.1136/amiajnl-2011-000464. https://europepmc.org/abstract/MED/21846786 amiajnl-2011-000464 - DOI - PMC - PubMed
1. Aronson AR. Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program. Proc AMIA Symp. 2001:17–21. https://europepmc.org/abstract/MED/11825149 D010001275 - PMC - PubMed
1. Yang X, Bian J, Fang R, Bjarnadottir RI, Hogan WR, Wu Y. Identifying relations of medications with adverse drug events using recurrent convolutional neural networks and gradient boosting. J Am Med Inform Assoc. 2020;27(1):65–72. doi: 10.1093/jamia/ocz144. https://europepmc.org/abstract/MED/31504605 5555856 - DOI - PMC - PubMed
1. Ahmed A, Abbasi A, Eickhoff C. Benchmarking modern named entity recognition techniques for free-text health record deidentification. AMIA Jt Summits Transl Sci Proc. 2021;2021:102–111. https://europepmc.org/abstract/MED/34457124 3477894 - PMC - PubMed
1. Uzuner O, Solti I, Cadag E. Extracting medication information from clinical text. J Am Med Inform Assoc. 2010;17(5):514–518. doi: 10.1136/jamia.2010.003947. https://europepmc.org/abstract/MED/20819854 17/5/514 - DOI - PMC - PubMed

Grants and funding

R01 GM141476/GM/NIGMS NIH HHS/United States

LinkOut - more resources

Full Text Sources

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study

Affiliations

Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

References

Grants and funding

LinkOut - more resources

Full Text Sources