. 2024 Mar 15;79(4):307-315.

doi: 10.1136/thorax-2023-220226.

Radiomics analysis to predict pulmonary nodule malignancy using machine learning approaches

Matthew T Warkentin^{1

2}, Hamad Al-Sawaihey¹, Stephen Lam^{3

4}, Geoffrey Liu^{2

5}, Brenda Diergaarde^{6

7}, Jian-Min Yuan^{7

8}, David O Wilson⁹, Sukhinder Atkar-Khattra⁴, Benjamin Grant⁵, Yonathan Brhane¹, Elham Khodayari-Moez¹, Kiera R Murison¹, Martin C Tammemagi¹⁰, Kieran R Campbell^{1

11}, Rayjean J Hung^{12

2}

Affiliations

¹ Prosserman Centre for Population Health Research, Lunenfeld-Tanenbaum Research Institute, Sinai Health, Toronto, Ontario, Canada.
² Dalla Lana School of Public Health, University of Toronto, Toronto, Ontario, Canada.
³ Department of Medicine, University of British Columbia, Vancouver, British Columbia, Canada.
⁴ Department of Integrative Oncology, British Columbia Cancer Research Institute, Vancouver, British Columbia, Canada.
⁵ Department of Medical Oncology and Hematology, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
⁶ Department of Human Genetics, University of Pittsburgh School of Public Health, Pittsburgh, Pennsylvania, USA.
⁷ Cancer Epidemiology and Prevention Program, UPMC Hillman Cancer Center, Pittsburgh, Pennsylvania, USA.
⁸ Department of Epidemiology, University of Pittsburgh School of Public Health, Pittsburgh, Pennsylvania, USA.
⁹ Department of Medicine, University of Pittsburgh, Pittsburgh, Pennsylvania, USA.
¹⁰ Cancer Control and Evidence Integration, Cancer Care Ontario, Toronto, Ontario, Canada.
¹¹ Department of Molecular Genetics, University of Toronto, Toronto, Ontario, Canada.
¹² Prosserman Centre for Population Health Research, Lunenfeld-Tanenbaum Research Institute, Sinai Health, Toronto, Ontario, Canada Rayjean.Hung@lunenfeld.ca.

PMID: 38195644
PMCID: PMC10947877
DOI: 10.1136/thorax-2023-220226

Radiomics analysis to predict pulmonary nodule malignancy using machine learning approaches

Matthew T Warkentin et al. Thorax. 2024.

. 2024 Mar 15;79(4):307-315.

doi: 10.1136/thorax-2023-220226.

Authors

Affiliations

¹ Prosserman Centre for Population Health Research, Lunenfeld-Tanenbaum Research Institute, Sinai Health, Toronto, Ontario, Canada.
² Dalla Lana School of Public Health, University of Toronto, Toronto, Ontario, Canada.
³ Department of Medicine, University of British Columbia, Vancouver, British Columbia, Canada.
⁴ Department of Integrative Oncology, British Columbia Cancer Research Institute, Vancouver, British Columbia, Canada.
⁵ Department of Medical Oncology and Hematology, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
⁶ Department of Human Genetics, University of Pittsburgh School of Public Health, Pittsburgh, Pennsylvania, USA.
⁷ Cancer Epidemiology and Prevention Program, UPMC Hillman Cancer Center, Pittsburgh, Pennsylvania, USA.
⁸ Department of Epidemiology, University of Pittsburgh School of Public Health, Pittsburgh, Pennsylvania, USA.
⁹ Department of Medicine, University of Pittsburgh, Pittsburgh, Pennsylvania, USA.
¹⁰ Cancer Control and Evidence Integration, Cancer Care Ontario, Toronto, Ontario, Canada.
¹¹ Department of Molecular Genetics, University of Toronto, Toronto, Ontario, Canada.
¹² Prosserman Centre for Population Health Research, Lunenfeld-Tanenbaum Research Institute, Sinai Health, Toronto, Ontario, Canada Rayjean.Hung@lunenfeld.ca.

PMID: 38195644
PMCID: PMC10947877
DOI: 10.1136/thorax-2023-220226

Abstract

Background: Low-dose CT screening can reduce lung cancer-related mortality. However, most screen-detected pulmonary abnormalities do not develop into cancer and it often remains challenging to identify malignant nodules, particularly among indeterminate nodules. We aimed to develop and assess prediction models based on radiological features to discriminate between benign and malignant pulmonary lesions detected on a baseline screen.

Methods: Using four international lung cancer screening studies, we extracted 2060 radiomic features for each of 16 797 nodules (513 malignant) among 6865 participants. After filtering out low-quality radiomic features, 642 radiomic and 9 epidemiological features remained for model development. We used cross-validation and grid search to assess three machine learning (ML) models (eXtreme Gradient Boosted Trees, random forest, least absolute shrinkage and selection operator (LASSO)) for their ability to accurately predict risk of malignancy for pulmonary nodules. We report model performance based on the area under the curve (AUC) and calibration metrics in the held-out test set.

Results: The LASSO model yielded the best predictive performance in cross-validation and was fit in the full training set based on optimised hyperparameters. Our radiomics model had a test-set AUC of 0.93 (95% CI 0.90 to 0.96) and outperformed the established Pan-Canadian Early Detection of Lung Cancer model (AUC 0.87, 95% CI 0.85 to 0.89) for nodule assessment. Our model performed well among both solid (AUC 0.93, 95% CI 0.89 to 0.97) and subsolid nodules (AUC 0.91, 95% CI 0.85 to 0.95).

Conclusions: We developed highly accurate ML models based on radiomic and epidemiological features from four international lung cancer screening studies that may be suitable for assessing indeterminate screen-detected pulmonary nodules for risk of malignancy.

Keywords: clinical epidemiology; imaging/CT MRI etc; lung cancer.

PubMed Disclaimer

Conflict of interest statement

Competing interests: None declared.

Figures

**Figure 1.**
Schematic for the analytic framework used in this study. Data were partitioned into training/validation and testing splits using group-based random sampling to ensure all nodules for a participant were in a single set to avoid data leakage. Radiomic features were extracted and subject to filtering to exclude low-quality and highly-redundant features. K-fold cross-validation was performed to identify the optimal machine learning (ML) model and the optimal set of hyperparameters. The final ML model was fitted to the entire training data set and tested for out-of-sample performance in the hold-out test data; discrimination and calibration performance metrics are reported.

**Figure 2.**
Receiver operating characteristic (ROC) curves for our INTEGRAL-Radiomics models and the established PanCan Model. Area under the curve (AUC) and 95% confidence intervals are reported.

**Figure 3.**
Calibration of our INTEGRAL-Radiomics model and the PanCan model (McWilliams *et al.*, 2013) in hold-out test-data. (A) Model-predicted risks versus observed risks across quintiles of model-predicted risks. The diagonal dashed line indicates perfect calibration. (B) Observed and expected (model-predicted) number of malignant pulmonary nodules (per 1,000 nodules) including the calibration ratio (Exp / Obs) and difference (Exp - Obs). Calibration ratios less than 1 (or differences less than 0) indicate underestimation of risk and ratios greater than 1 (or differences greater than 0) indicate overestimation of risk.

See this image and copyright information in PMC

References

1. Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: a cancer journal for clinicians. 2021; - PubMed
1. Howlader N, Noone A, Krapcho M, Miller D, Bishop K, Kosary C, Yu M, Ruhl J, Tatalovich Z, Mariotto A, others. SEER cancer statistics review, 1975–2014, national cancer institute. Bethesda, MD. 2017;1–12.
1. Koning HJ de, Aalst CM van der, Jong PA de, Scholten ET, Nackaerts, Heuvelmans MA, Lammers J-WJ, Weenink C, Yousaf-Khan U, Horeweg N, others. Reduced lung-cancer mortality with volume CT screening in a randomized trial. New England Journal of Medicine. 2020;382(6):503–13. - PubMed
1. Team NLSTR. Reduced lung-cancer mortality with low-dose computed tomographic screening. New England Journal of Medicine. 2011;365(5):395–409. - PMC - PubMed
1. National Lung Screening Trial Research Team. Lung cancer incidence and mortality with extended follow-up in the national lung screening trial. Journal of Thoracic Oncology. 2019;14(10):1732–42. - PMC - PubMed

Publication types

Actions
Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

Grants and funding

U19 CA203654/CA/NCI NIH HHS/United States

LinkOut - more resources

Full Text Sources
Medical
- MedlinePlus Health Information

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Radiomics analysis to predict pulmonary nodule malignancy using machine learning approaches

Affiliations

Radiomics analysis to predict pulmonary nodule malignancy using machine learning approaches

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

References

Publication types

MeSH terms

Grants and funding

LinkOut - more resources

Full Text Sources

Medical