. 2024 Mar 12;24(1):72.

doi: 10.1186/s12911-024-02459-6.

Assessing the research landscape and clinical utility of large language models: a scoping review

Ye-Jean Park¹, Abhinav Pillai^#², Jiawen Deng^#³, Eddie Guo², Mehul Gupta², Mike Paget², Christopher Naugler²

Affiliations

¹ Temerty Faculty of Medicine, University of Toronto, 1 King's College Cir, M5S 1A8, Toronto, ON, Canada. yejean.park@mail.utoronto.ca.
² Cumming School of Medicine, University of Calgary, 3330 Hospital Dr NW, T2N 4N1, Calgary, AB, Canada.
³ Temerty Faculty of Medicine, University of Toronto, 1 King's College Cir, M5S 1A8, Toronto, ON, Canada.

^# Contributed equally.

PMID: 38475802
PMCID: PMC10936025
DOI: 10.1186/s12911-024-02459-6

Assessing the research landscape and clinical utility of large language models: a scoping review

Ye-Jean Park et al. BMC Med Inform Decis Mak. 2024.

. 2024 Mar 12;24(1):72.

doi: 10.1186/s12911-024-02459-6.

Authors

Ye-Jean Park¹, Abhinav Pillai^#², Jiawen Deng^#³, Eddie Guo², Mehul Gupta², Mike Paget², Christopher Naugler²

Affiliations

¹ Temerty Faculty of Medicine, University of Toronto, 1 King's College Cir, M5S 1A8, Toronto, ON, Canada. yejean.park@mail.utoronto.ca.
² Cumming School of Medicine, University of Calgary, 3330 Hospital Dr NW, T2N 4N1, Calgary, AB, Canada.
³ Temerty Faculty of Medicine, University of Toronto, 1 King's College Cir, M5S 1A8, Toronto, ON, Canada.

^# Contributed equally.

PMID: 38475802
PMCID: PMC10936025
DOI: 10.1186/s12911-024-02459-6

Abstract
in English, German

Importance: Large language models (LLMs) like OpenAI's ChatGPT are powerful generative systems that rapidly synthesize natural language responses. Research on LLMs has revealed their potential and pitfalls, especially in clinical settings. However, the evolving landscape of LLM research in medicine has left several gaps regarding their evaluation, application, and evidence base.

Objective: This scoping review aims to (1) summarize current research evidence on the accuracy and efficacy of LLMs in medical applications, (2) discuss the ethical, legal, logistical, and socioeconomic implications of LLM use in clinical settings, (3) explore barriers and facilitators to LLM implementation in healthcare, (4) propose a standardized evaluation framework for assessing LLMs' clinical utility, and (5) identify evidence gaps and propose future research directions for LLMs in clinical applications.

Evidence review: We screened 4,036 records from MEDLINE, EMBASE, CINAHL, medRxiv, bioRxiv, and arXiv from January 2023 (inception of the search) to June 26, 2023 for English-language papers and analyzed findings from 55 worldwide studies. Quality of evidence was reported based on the Oxford Centre for Evidence-based Medicine recommendations.

Findings: Our results demonstrate that LLMs show promise in compiling patient notes, assisting patients in navigating the healthcare system, and to some extent, supporting clinical decision-making when combined with human oversight. However, their utilization is limited by biases in training data that may harm patients, the generation of inaccurate but convincing information, and ethical, legal, socioeconomic, and privacy concerns. We also identified a lack of standardized methods for evaluating LLMs' effectiveness and feasibility.

Conclusions and relevance: This review thus highlights potential future directions and questions to address these limitations and to further explore LLMs' potential in enhancing healthcare delivery.

Question What is the current state of Large Language Models’ (LLMs) application in clinical settings, and what are the primary challenges and opportunities associated with their integration? Findings This scoping review, analyzing 55 studies, indicates that while LLMs, including OpenAI’s ChatGPT, show potential in compiling patient notes, aiding in healthcare navigation, and supporting clinical decision-making, their use is constrained by data biases, the generation of plausible but incorrect information, and various ethical and privacy concerns. A significant variability in the rigor of studies, especially in evaluating LLM responses, calls for standardized evaluation methods, including established metrics like ROUGE, METEOR, G-Eval, and MultiMedQA. Meaning The findings suggest a need for enhanced methodologies in LLM research, stressing the importance of integrating real patient data and considering social determinants of health, to improve the applicability and safety of LLMs in clinical environments.

Keywords: ChatGPT; Clinical settings; Large language models; Natural language processing; Scoping review.

PubMed Disclaimer

Conflict of interest statement

The authors declare no competing interests.

Figures

**Fig. 1**
Flowchart of the search strategy

**Fig. 2**
Number of articles published over the timespan of January 2022 to June 2023

**Fig. 3**
Types of included studies (n = 55). Preprints were the most common (n = 21) whereas case reports were the least common (n = 2)

See this image and copyright information in PMC

References

1. Yang X, Chen A, PourNejatian N, Shin HC, Smith KE, Parisien C, et al. A large language model for electronic health records. NPJ Digit Med. 2022;5(1):194. doi: 10.1038/s41746-022-00742-2. - DOI - PMC - PubMed
1. OpenAI. Introducing ChatGPT [Internet]. [cited 2023 May 2]. Available from: https://openai.com/blog/chatgpt.
1. Devlin J, Chang MW, Lee K, Toutanova K, BERT. Pre-training of deep bidirectional Transformers for language understanding [Internet]. arXiv. 2018. Available from: https://arxiv.org/abs/1810.04805.
1. Levine DM, Tuwani R, Kompa B, Varma A, Finlayson SG, Mehrotra A et al. The Diagnostic and Triage Accuracy of the GPT-3 Artificial Intelligence Model [Internet]. medRxiv. 2023. 10.1101/2023.01.30.23285067. - PubMed
1. Stewart J, Lu J, Goudie A, Arendts G, Meka SA, Freeman S et al. Applications of natural language processing at emergency department triage: A systematic review [Internet]. bioRxiv. 2022. 10.1101/2022.12.20.22283735. - PMC - PubMed

Publication types

Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Assessing the research landscape and clinical utility of large language models: a scoping review

Affiliations

Assessing the research landscape and clinical utility of large language models: a scoping review

Authors

Affiliations

Abstract
in English, German

Conflict of interest statement

Figures

References

Publication types

MeSH terms

LinkOut - more resources

Full Text Sources

Abstract in English, German

Conflict of interest statement

Figures

References

Publication types

MeSH terms

LinkOut - more resources

Full Text Sources

Abstract
in English, German