Deep Learning Accurately Predicts Estrogen Receptor Status in Breast Cancer Metabolomics Data
- PMID: 29110491
- PMCID: PMC5759031
- DOI: 10.1021/acs.jproteome.7b00595
Deep Learning Accurately Predicts Estrogen Receptor Status in Breast Cancer Metabolomics Data
Abstract
Metabolomics holds the promise as a new technology to diagnose highly heterogeneous diseases. Conventionally, metabolomics data analysis for diagnosis is done using various statistical and machine learning based classification methods. However, it remains unknown if deep neural network, a class of increasingly popular machine learning methods, is suitable to classify metabolomics data. Here we use a cohort of 271 breast cancer tissues, 204 positive estrogen receptor (ER+), and 67 negative estrogen receptor (ER-) to test the accuracies of feed-forward networks, a deep learning (DL) framework, as well as six widely used machine learning models, namely random forest (RF), support vector machines (SVM), recursive partitioning and regression trees (RPART), linear discriminant analysis (LDA), prediction analysis for microarrays (PAM), and generalized boosted models (GBM). DL framework has the highest area under the curve (AUC) of 0.93 in classifying ER+/ER- patients, compared to the other six machine learning algorithms. Furthermore, the biological interpretation of the first hidden layer reveals eight commonly enriched significant metabolomics pathways (adjusted P-value <0.05) that cannot be discovered by other machine learning methods. Among them, protein digestion and absorption and ATP-binding cassette (ABC) transporters pathways are also confirmed in integrated analysis between metabolomics and gene expression data in these samples. In summary, deep learning method shows advantages for metabolomics based breast cancer ER status classification, with both the highest prediction accuracy (AUC = 0.93) and better revelation of disease biology. We encourage the adoption of feed-forward networks based deep learning method in the metabolomics research community for classification.
Keywords: bioinformatics; breast cancer; deep learning; estrogen receptor; metabolomics.
Conflict of interest statement
The authors declare no competing financial interest.
Figures
References
-
- Breast Cancer: Prevention and Control; World Health Organization, 2017. http://www.who.int/cancer/detection/breastcancer/en/index1.html (accessed October 10, 2017).
-
- About Breast Cancer; American Cancer Society, 2017. https://www.cancer.org/cancer/breast-cancer/about/how-common-is-breast-c... (accessed September 21, 2017).
-
- Carey L. A.; Perou C. M.; Livasy C. A.; Dressler L. G.; Cowan D.; Conway K.; Karaca G.; Troester M. A.; Tse C. K.; Edmiston S.; Deming S. L.; Geradts J.; Cheang M. C.; Nielsen T. O.; Moorman P. G.; Earp H. S.; Millikan R. C. Race, breast cancer subtypes, and survival in the Carolina Breast Cancer Study. JAMA 2006, 295 (21), 2492–2502. 10.1001/jama.295.21.2492. - DOI - PubMed
-
- O’Brien K. M.; Cole S. R.; Tse C. K.; Perou C. M.; Carey L. A.; Foulkes W. D.; Dressler L. G.; Geradts J.; Millikan R. C. Intrinsic breast tumor subtypes, race, and long-term survival in the Carolina Breast Cancer Study. Clin. Cancer Res. 2010, 16 (24), 6100–6110. 10.1158/1078-0432.CCR-10-1533. - DOI - PMC - PubMed
-
- Haque R.; Ahmed S. A.; Inzhakova G.; Shi J.; Avila C.; Polikoff J.; Bernstein L.; Enger S. M.; Press M. F. Impact of breast cancer subtypes and treatment on survival: an analysis spanning two decades. Cancer Epidemiol., Biomarkers Prev. 2012, 21 (10), 1848–1855. 10.1158/1055-9965.EPI-12-0474. - DOI - PMC - PubMed
Publication types
MeSH terms
Substances
Grants and funding
LinkOut - more resources
Full Text Sources
Other Literature Sources
Medical
Miscellaneous
