Automated Extraction of Tumor Staging and Diagnosis Information From Surgical Pathology Reports

Sajjad Abedian¹, Evan T Sholle^{1

2}, Prakash M Adekkanattu¹, Marika M Cusick¹, Stephanie E Weiner¹, Jonathan E Shoag², Jim C Hu², Thomas R Campion Jr^{1

3

4

5}

Affiliations

¹ Information Technologies and Services Department, Weill Cornell Medicine, New York, NY.
² Department of Population Health Sciences, Weill Cornell Medicine, New York, NY.
³ Department of Urology, Weill Cornell Medicine, New York, NY.
⁴ Clinical and Translational Science Center, Weill Cornell Medicine, New York, NY.
⁵ Department of Pediatrics, Weill Cornell Medicine, New York, NY.

PMID: 34694896
PMCID: PMC8812635
DOI: 10.1200/CCI.21.00065

Automated Extraction of Tumor Staging and Diagnosis Information From Surgical Pathology Reports

Sajjad Abedian et al. JCO Clin Cancer Inform. 2021 Oct.

. 2021 Oct:5:1054-1061.

doi: 10.1200/CCI.21.00065.

Authors

Sajjad Abedian¹, Evan T Sholle^{1

2}, Prakash M Adekkanattu¹, Marika M Cusick¹, Stephanie E Weiner¹, Jonathan E Shoag², Jim C Hu², Thomas R Campion Jr^{1

3

4

5}

Affiliations

¹ Information Technologies and Services Department, Weill Cornell Medicine, New York, NY.
² Department of Population Health Sciences, Weill Cornell Medicine, New York, NY.
³ Department of Urology, Weill Cornell Medicine, New York, NY.
⁴ Clinical and Translational Science Center, Weill Cornell Medicine, New York, NY.
⁵ Department of Pediatrics, Weill Cornell Medicine, New York, NY.

PMID: 34694896
PMCID: PMC8812635
DOI: 10.1200/CCI.21.00065

Abstract

Purpose: Typically stored as unstructured notes, surgical pathology reports contain data elements valuable to cancer research that require labor-intensive manual extraction. Although studies have described natural language processing (NLP) of surgical pathology reports to automate information extraction, efforts have focused on specific cancer subtypes rather than across multiple oncologic domains. To address this gap, we developed and evaluated an NLP method to extract tumor staging and diagnosis information across multiple cancer subtypes.

Methods: The NLP pipeline was implemented on an open-source framework called Leo. We used a total of 555,681 surgical pathology reports of 329,076 patients to develop the pipeline and evaluated our approach on subsets of reports from patients with breast, prostate, colorectal, and randomly selected cancer subtypes.

Results: Averaged across all four cancer subtypes, the NLP pipeline achieved an accuracy of 1.00 for International Classification of Diseases, Tenth Revision codes, 0.89 for T staging, 0.90 for N staging, and 0.97 for M staging. It achieved an F1 score of 1.00 for International Classification of Diseases, Tenth Revision codes, 0.88 for T staging, 0.90 for N staging, and 0.24 for M staging.

Conclusion: The NLP pipeline was developed to extract tumor staging and diagnosis information across multiple cancer subtypes to support the research enterprise in our institution. Although it was not possible to demonstrate generalizability of our NLP pipeline to other institutions, other institutions may find value in adopting a similar NLP approach-and reusing code available at GitHub-to support the oncology research enterprise with elements extracted from surgical pathology reports.

PubMed Disclaimer

Conflict of interest statement

Evan T. SholleStock and Other Ownership Interests: Moderna Therapeutics Jonathan E. ShoagResearch Funding: Bristol Myers Squibb Foundation Jim C. HuSpeakers' Bureau: Genomic Health, Intuitive SurgicalTravel, Accommodations, Expenses: Intuitive SurgicalNo other potential conflicts of interest were reported.

Figures

**FIG 1.**
Mock-up pathology report that contains sections with pathologic staging and ICD-10 code information. h/o, history of ICD-10, International Classification of Diseases, Tenth Revision; PSA, prostate-specific antigen.

**FIG 2.**
(A) NLP logic implemented for extracting TNM staging and diagnosis codes from surgical pathology reports. (B) NLP logic implementation on a pathology report to extract TNM staging from a surgical pathology report. NLP, natural language processing; pTNM, Pathological Tumor-Node-Metastasis.

See this image and copyright information in PMC

References

1. Griffon N, Charlet J, Darmoni SJ.Managing free text for secondary use of health data Yearb Med Inform 9167–1692014 - PMC - PubMed
1. Yim W-W, Yetisgen M, Harris WP, et al. Natural language processing in oncology: A review JAMA Oncol 2797–8042016 - PubMed
1. Cancer Protocol Templates|College of American Pathologists. https://www.cap.org/protocols-and-guidelines/cancer-reporting-tools/canc...
1. Renshaw AA, Mena-Allauca M, Gould EW, et al. Synoptic reporting: Evidence-based review and future directions JCO Clin Cancer Inform 21–92018 - PMC - PubMed
1. Kim BJ, Merchant M, Zheng C, et al. A natural language processing program effectively extracts key pathologic findings from radical prostatectomy reports J Endourol 281474–14782014 - PubMed

Publication types

Actions
Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions

Grants and funding

UL1 TR000457/TR/NCATS NIH HHS/United States

LinkOut - more resources

Full Text Sources
Miscellaneous
- NCI CPTAC Assay Portal

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Automated Extraction of Tumor Staging and Diagnosis Information From Surgical Pathology Reports

Affiliations

Automated Extraction of Tumor Staging and Diagnosis Information From Surgical Pathology Reports

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

References

Publication types

MeSH terms

Grants and funding

LinkOut - more resources

Full Text Sources

Miscellaneous