Large language models can accurately populate Vascular Quality Initiative procedural databases using narrative operative reports

Colleen P Flanagan¹, Karen Trang², Joyce Nacario³, Peter A Schneider³, Warren J Gasper³, Michael S Conte³, Elizabeth C Wick², Allan M Conway³

Affiliations

¹ Division of Vascular and Endovascular Surgery, Department of Surgery, University of California San Francisco, San Francisco, CA; Division of Clinical Informatics and Digital Transformation, Department of Medicine, University of California San Francisco, San Francisco, CA. Electronic address: Colleen.Flanagan@ucsf.edu.
² Division of Clinical Informatics and Digital Transformation, Department of Medicine, University of California San Francisco, San Francisco, CA; Division of General Surgery, Department of Surgery, University of California San Francisco, San Francisco, CA.
³ Division of Vascular and Endovascular Surgery, Department of Surgery, University of California San Francisco, San Francisco, CA.

PMID: 39694151
DOI: 10.1016/j.jvs.2024.12.002

Free article

Large language models can accurately populate Vascular Quality Initiative procedural databases using narrative operative reports

Colleen P Flanagan et al. J Vasc Surg. 2025 Apr.

Free article

. 2025 Apr;81(4):973-982.

doi: 10.1016/j.jvs.2024.12.002. Epub 2024 Dec 16.

Authors

Colleen P Flanagan¹, Karen Trang², Joyce Nacario³, Peter A Schneider³, Warren J Gasper³, Michael S Conte³, Elizabeth C Wick², Allan M Conway³

Affiliations

¹ Division of Vascular and Endovascular Surgery, Department of Surgery, University of California San Francisco, San Francisco, CA; Division of Clinical Informatics and Digital Transformation, Department of Medicine, University of California San Francisco, San Francisco, CA. Electronic address: Colleen.Flanagan@ucsf.edu.
² Division of Clinical Informatics and Digital Transformation, Department of Medicine, University of California San Francisco, San Francisco, CA; Division of General Surgery, Department of Surgery, University of California San Francisco, San Francisco, CA.
³ Division of Vascular and Endovascular Surgery, Department of Surgery, University of California San Francisco, San Francisco, CA.

PMID: 39694151
DOI: 10.1016/j.jvs.2024.12.002

Abstract

Objective: Participation in the Vascular Quality Initiative (VQI) provides important resources to surgeons, but the ability to do so is often limited by time and data entry personnel. Large language models (LLMs) such as ChatGPT (OpenAI) are examples of generative artificial intelligence products that may help bridge this gap. Trained on large volumes of data, the models are used for natural language processing and text generation. We evaluated the ability of LLMs to accurately populate VQI procedural databases using operative reports.

Methods: A single-center, retrospective study was performed using institutional VQI data from 2021 to 2023. The most recent procedures for carotid endarterectomy (CEA), endovascular aneurysm repair (EVAR), and infrainguinal lower extremity bypass (LEB) were analyzed using Versa, a HIPAA (Health Insurance Portability and Accountability Act)-compliant institutional version of ChatGPT. We created an automated function to analyze operative reports and generate a shareable VQI file using two models: gpt-35-turbo and gpt-4. Application of the LLMs was accomplished with a cloud-based programming interface. The outputs of this model were compared with VQI data for accuracy. We defined a metric as "unavailable" to the LLM if it was discussed by surgeons in <20% of operative reports.

Results: A total of 150 operative notes were analyzed, including 50 CEA, 50 EVAR, and 50 LEB. These procedural VQI databases included 25, 179, and 51 metrics, respectively. For all fields, gpt-35-turbo had a median accuracy of 84.0% for CEA (interquartile range [IQR]: 80.0%-88.0%), 92.2% for EVAR (IQR: 87.2%-94.0%), and 84.3% for LEB (IQR: 80.2%-88.1%). A total of 3 of 25, 6 of 179, and 7 of 51 VQI variables were unavailable in the operative reports, respectively. Excluding metric information routinely unavailable in operative reports, the median accuracy rate was 95.5% for each CEA procedure (IQR: 90.9%-100.0%), 94.8% for EVAR (IQR: 92.2%-98.5%), and 93.2% for LEB (IQR: 90.2%-96.4%). Across procedures, gpt-4 did not meaningfully improve performance compared with gpt-35 (P = .97, .85, and .95 for CEA, EVAR, and LEB overall performance, respectively). The cost for 150 operative reports analyzed with gpt-35-turbo and gpt-4 was $0.12 and $3.39, respectively.

Conclusions: LLMs can accurately populate VQI procedural databases with both structured and unstructured data, while incurring only minor processing costs. Increased workflow efficiency may improve center ability to successfully participate in the VQI. Further work examining other VQI databases and methods to increase accuracy is needed.

Keywords: Generative artificial intelligence; Large language models; Quality reporting.

PubMed Disclaimer

Conflict of interest statement

Disclosures None.

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
Miscellaneous
- NCI CPTAC Assay Portal

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Large language models can accurately populate Vascular Quality Initiative procedural databases using narrative operative reports

Affiliations

Large language models can accurately populate Vascular Quality Initiative procedural databases using narrative operative reports

Authors

Affiliations

Abstract

Conflict of interest statement

MeSH terms

LinkOut - more resources

Full Text Sources

Miscellaneous