Almanac - Retrieval-Augmented Language Models for Clinical Medicine

Cyril Zakka¹, Rohan Shad², Akash Chaurasia³, Alex R Dalal¹, Jennifer L Kim¹, Michael Moor³, Robyn Fong³, Curran Phillips¹, Kevin Alexander⁴, Euan Ashley⁴, Jack Boyd¹, Kathleen Boyd⁵, Karen Hirsch⁶, Curt Langlotz⁷, Rita Lee¹, Joanna Melia⁸, Joanna Nelson⁹, Karim Sallam⁴, Stacey Tullis¹, Melissa Ann Vogelsong¹⁰, John Patrick Cunningham¹¹, William Hiesinger¹

Affiliations

¹ Department of Cardiothoracic Surgery, Stanford Medicine, Stanford, CA.
² Division of Cardiovascular Surgery, Penn Medicine, Philadelphia.
³ Department of Computer Science, Stanford University, Stanford, CA.
⁴ Division of Cardiovascular Medicine, Stanford Medicine, Stanford, CA.
⁵ Department of Pediatrics, Stanford Medicine, Stanford, CA.
⁶ Department of Neurology, Stanford Medicine, Stanford, CA.
⁷ Department of Radiology and Biomedical Informatics, Stanford Medicine, Stanford, CA.
⁸ Division of Gastroenterology and Hepatology, Johns Hopkins Medicine, Baltimore.
⁹ Division of Infectious Diseases, Stanford Medicine, Stanford, CA.
¹⁰ Division of Anesthesia, Stanford Medicine, Stanford, CA.
¹¹ Department of Statistics, Columbia University, New York.

PMID: 38343631
PMCID: PMC10857783
DOI: 10.1056/aioa2300068

Almanac - Retrieval-Augmented Language Models for Clinical Medicine

Cyril Zakka et al. NEJM AI. 2024 Feb.

. 2024 Feb;1(2):10.1056/aioa2300068.

doi: 10.1056/aioa2300068. Epub 2024 Jan 25.

Authors

Affiliations

¹ Department of Cardiothoracic Surgery, Stanford Medicine, Stanford, CA.
² Division of Cardiovascular Surgery, Penn Medicine, Philadelphia.
³ Department of Computer Science, Stanford University, Stanford, CA.
⁴ Division of Cardiovascular Medicine, Stanford Medicine, Stanford, CA.
⁵ Department of Pediatrics, Stanford Medicine, Stanford, CA.
⁶ Department of Neurology, Stanford Medicine, Stanford, CA.
⁷ Department of Radiology and Biomedical Informatics, Stanford Medicine, Stanford, CA.
⁸ Division of Gastroenterology and Hepatology, Johns Hopkins Medicine, Baltimore.
⁹ Division of Infectious Diseases, Stanford Medicine, Stanford, CA.
¹⁰ Division of Anesthesia, Stanford Medicine, Stanford, CA.
¹¹ Department of Statistics, Columbia University, New York.

PMID: 38343631
PMCID: PMC10857783
DOI: 10.1056/aioa2300068

Abstract

Background: Large language models (LLMs) have recently shown impressive zero-shot capabilities, whereby they can use auxiliary data, without the availability of task-specific training examples, to complete a variety of natural language tasks, such as summarization, dialogue generation, and question answering. However, despite many promising applications of LLMs in clinical medicine, adoption of these models has been limited by their tendency to generate incorrect and sometimes even harmful statements.

Methods: We tasked a panel of eight board-certified clinicians and two health care practitioners with evaluating Almanac, an LLM framework augmented with retrieval capabilities from curated medical resources for medical guideline and treatment recommendations. The panel compared responses from Almanac and standard LLMs (ChatGPT-4, Bing, and Bard) versus a novel data set of 314 clinical questions spanning nine medical specialties.

Results: Almanac showed a significant improvement in performance compared with the standard LLMs across axes of factuality, completeness, user preference, and adversarial safety.

Conclusions: Our results show the potential for LLMs with access to domain-specific corpora to be effective in clinical decision-making. The findings also underscore the importance of carefully testing LLMs before deployment to mitigate their shortcomings. (Funded by the National Institutes of Health, National Heart, Lung, and Blood Institute.).

PubMed Disclaimer

Figures

**Figure 1.. Almanac Overview.**
When presented with a query, Almanac uses external tools to retrieve relevant information before synthesizing a response with citations referencing source material. With this framework, large language model (LLM) outputs remain grounded in truth while providing a reliable way of fact-checking.

**Figure 2.. Heat Maps of the Nemenyi P Values for Factuality, Completeness, and Preference for Model Pairs across ClinicalQA.**
Red denotes significant differences at P<0.01; blue denotes nonsignificant differences.

See this image and copyright information in PMC

Update of

Almanac: Retrieval-Augmented Language Models for Clinical Medicine.
Zakka C, Chaurasia A, Shad R, Dalal AR, Kim JL, Moor M, Alexander K, Ashley E, Boyd J, Boyd K, Hirsch K, Langlotz C, Nelson J, Hiesinger W. Zakka C, et al. Res Sq [Preprint]. 2023 May 2:rs.3.rs-2883198. doi: 10.21203/rs.3.rs-2883198/v1. Res Sq. 2023. Update in: NEJM AI. 2024 Feb;1(2). doi: 10.1056/aioa2300068. PMID: 37205549 Free PMC article. Updated. Preprint.

References

1. Brown TB, Mann B, Ryder N, et al. Language models are few-shot learners. July 22, 2020. (https://arxiv.org/abs/2005.14165). Preprint.
1. Chen M, Tworek J, Jun H, et al. Evaluating large language models trained on code. July 14, 2021. (https://arxiv.org/abs/2107.03374). Preprint.
1. Wei C, Xie SM, Ma T. Why do pretrained language models help in downstream tasks? An analysis of head and prompt tuning. June 2021. (https://arxiv.org/abs/2106.09226). Preprint.
1. Devlin J, Chang M-W, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. October 2018. (https://arxiv.org/abs/1810.04805). Preprint.
1. Wei J, Tay Y, Bommasani R, et al. Emergent abilities of large language models. June 2022. (https://arxiv.org/abs/2206.07682). Preprint.

Grants and funding

R01 HL157235/HL/NHLBI NIH HHS/United States

LinkOut - more resources

Full Text Sources
- Europe PubMed Central
- PubMed Central
Other Literature Sources
- The Lens - Patent Citations Database

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Almanac - Retrieval-Augmented Language Models for Clinical Medicine

Affiliations

Almanac - Retrieval-Augmented Language Models for Clinical Medicine

Authors

Affiliations

Abstract

Figures

Update of

References

Grants and funding

LinkOut - more resources

Full Text Sources

Other Literature Sources