Developing prompts from large language model for extracting clinical information from pathology and ultrasound reports in breast cancer

Hyeon Seok Choi¹, Jun Yeong Song¹, Kyung Hwan Shin^{1

2}, Ji Hyun Chang¹, Bum-Sup Jang¹

Affiliations

¹ Department of Radiation Oncology, Seoul National University Hospital, Seoul National University College of Medicine, Seoul, Korea.
² Institute of Radiation Medicine, Seoul National University Medical Research Center, Seoul, Korea.

PMID: 37793630
PMCID: PMC10556835
DOI: 10.3857/roj.2023.00633

Developing prompts from large language model for extracting clinical information from pathology and ultrasound reports in breast cancer

Hyeon Seok Choi et al. Radiat Oncol J. 2023 Sep.

. 2023 Sep;41(3):209-216.

doi: 10.3857/roj.2023.00633. Epub 2023 Sep 21.

Authors

Hyeon Seok Choi¹, Jun Yeong Song¹, Kyung Hwan Shin^{1

2}, Ji Hyun Chang¹, Bum-Sup Jang¹

Affiliations

¹ Department of Radiation Oncology, Seoul National University Hospital, Seoul National University College of Medicine, Seoul, Korea.
² Institute of Radiation Medicine, Seoul National University Medical Research Center, Seoul, Korea.

PMID: 37793630
PMCID: PMC10556835
DOI: 10.3857/roj.2023.00633

Abstract

Purpose: We aimed to evaluate the time and cost of developing prompts using large language model (LLM), tailored to extract clinical factors in breast cancer patients and their accuracy.

Materials and methods: We collected data from reports of surgical pathology and ultrasound from breast cancer patients who underwent radiotherapy from 2020 to 2022. We extracted the information using the Generative Pre-trained Transformer (GPT) for Sheets and Docs extension plugin and termed this the "LLM" method. The time and cost of developing the prompts with LLM methods were assessed and compared with those spent on collecting information with "full manual" and "LLM-assisted manual" methods. To assess accuracy, 340 patients were randomly selected, and the extracted information by LLM method were compared with those collected by "full manual" method.

Results: Data from 2,931 patients were collected. We developed 12 prompts for Extract function and 12 for Format function to extract and standardize the information. The overall accuracy was 87.7%. For lymphovascular invasion, it was 98.2%. Developing and processing the prompts took 3.5 hours and 15 minutes, respectively. Utilizing the ChatGPT application programming interface cost US $65.8 and when factoring in the estimated wage, the total cost was US $95.4. In an estimated comparison, "LLM-assisted manual" and "LLM" methods were time- and cost-efficient compared to the "full manual" method.

Conclusion: Developing and facilitating prompts for LLM to derive clinical factors was efficient to extract crucial information from huge medical records. This study demonstrated the potential of the application of natural language processing using LLM model in breast cancer patients. Prompts from the current study can be re-used for other research to collect clinical information.

Keywords: Ai artificial intelligence; Automatic data processing; Breast cancer; Clinical reports; Natural language processing.

PubMed Disclaimer

Conflict of interest statement

Conflict of Interest

No potential conflict of interest relevant to this article was reported.

Figures

**Fig. 1.**
The schema of the current study. Using ChatGPT 3.5 model, information about clinical T and N stage was extracted from ultrasound readings, and pathologic T and N stage and additional factors were extracted from pathology readings. Then, trimming was performed and organized in tabular form. For validation, a sample was randomly selected to evaluate the accuracy. GPT, Generative Pre-trained Transformer; IHC, immunohistochemistry.

See this image and copyright information in PMC

References

1. National Comprehensive Cancer Network NCCN Clinical Practice Guidelines in Oncology: breast cancer [Internet]. Plymouth Meeting, PA: National Comprehensive Cancer Network; 2023 [cited 2023 Sep 13]. Available from: https://www.nccn.org/professionals/physician_gls/pdf/breast.pdf.
1. Park HJ, Shin KH, Kim JH, et al. Incorporating risk factors to identify the indication of post-mastectomy radiotherapy in N1 breast cancer treated with optimal systemic therapy: a multicenter analysis in Korea (KROG 14-23) Cancer Res Treat. 2017;49:739–47. - PMC - PubMed
1. Yamada A, Hayashi N, Kumamaru H, et al. Prognostic impact of postoperative radiotherapy in patients with breast cancer and with pT1-2 and 1-3 lymph node metastases: a retrospective cohort study based on the Japanese Breast Cancer Registry. Eur J Cancer. 2022;172:31–40. - PubMed
1. Jwa E, Shin KH, Lim HW, et al. Identification of risk factors for locoregional recurrence in breast cancer patients with nodal stage N0 and N1: who could benefit from post-mastectomy radiotherapy? PLoS One. 2015;10:e0145463. - PMC - PubMed
1. Viani GA, Godoi da Silva LB, Viana BS. Patients with N1 breast cancer: who could benefit from supraclavicular fossa radiotherapy? Breast. 2014;23:749–53. - PubMed

Grants and funding

LinkOut - more resources

Full Text Sources

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Developing prompts from large language model for extracting clinical information from pathology and ultrasound reports in breast cancer

Affiliations

Developing prompts from large language model for extracting clinical information from pathology and ultrasound reports in breast cancer

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

References

Grants and funding

LinkOut - more resources

Full Text Sources