Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
Review
. 2025 May;69(5):e374-e394.
doi: 10.1016/j.jpainsymman.2025.01.019. Epub 2025 Jan 31.

NLP for Analyzing Electronic Health Records and Clinical Notes in Cancer Research: A Review

Affiliations
Review

NLP for Analyzing Electronic Health Records and Clinical Notes in Cancer Research: A Review

Muhammad Bilal et al. J Pain Symptom Manage. 2025 May.

Abstract

This review examines the application of natural language processing (NLP) techniques in cancer research using electronic health records (EHRs) and clinical notes. It addresses gaps in existing literature by providing a broader perspective than previous studies focused on specific cancer types or applications. A comprehensive literature search in the Scopus database identified 94 relevant studies published between 2019 and 2024. The analysis revealed a growing trend in NLP applications for cancer research, with information extraction (47 studies) and text classification (40 studies) emerging as predominant NLP tasks, followed by named entity recognition (7 studies). Among cancer types, breast, lung, and colorectal cancers were found to be the most studied. A significant shift from rule-based and traditional machine learning approaches to advanced deep learning techniques and transformer-based models was observed. It was found that dataset sizes used in existing studies varied widely, ranging from small, manually annotated datasets to large-scale EHRs. The review highlighted key challenges, including the limited generalizability of proposed solutions and the need for improved integration into clinical workflows. While NLP techniques show significant potential in analyzing EHRs and clinical notes for cancer research, future work should focus on improving model generalizability, enhancing robustness in handling complex clinical language, and expanding applications to understudied cancer types. The integration of NLP tools into palliative medicine and addressing ethical considerations remain crucial for utilizing the full potential of NLP in enhancing cancer diagnosis, treatment, and patient outcomes. This review provides valuable insights into the current state and future directions of NLP applications in cancer research.

Keywords: Cancer; Clinical notes; Electronic health records; Information extraction; Natural language processing; Text classification.

PubMed Disclaimer

Conflict of interest statement

Disclosures and Acknowledgments The authors have declared no conflict of interest. This research received no specific funding/grant from any funding agency in the public, commercial, or not-for-profit sectors.

Similar articles

LinkOut - more resources