Machine learning with the TCGA-HNSC dataset: improving usability by addressing inconsistency, sparsity, and high-dimensionality

Michael C Rendleman¹, John M Buatti², Terry A Braun³, Brian J Smith⁴, Chibuzo Nwakama⁵, Reinhard R Beichel^{6

7}, Bart Brown⁸, Thomas L Casavant^{5

3}

Affiliations

¹ Department of Electrical and Computer Engineering, Center for Bioinformatics and Computational Biology, University of Iowa, 5017 Seamans Center, Iowa City, IA, 52242, USA. michael-rendleman@uiowa.edu.
² Department of Radiation Oncology, Carver College of Medicine, University of Iowa Carver College of Medicine, LL-W Pomerantz Family Pavilion, 200 Hawkins Drive, Iowa City, IA, 52242-1089, USA.
³ Department of Biomedical Engineering, Center for Bioinformatics and Computational Biology, University of Iowa, 5017 Seamans Center, Iowa City, IA, 52242, USA.
⁴ Department of Biostatistics, University of Iowa, 145 N. Riverside Drive, 100 CPHB, Iowa City, IA, 52242, USA.
⁵ Department of Electrical and Computer Engineering, Center for Bioinformatics and Computational Biology, University of Iowa, 5017 Seamans Center, Iowa City, IA, 52242, USA.
⁶ Iowa Institute for Biomedical Imaging, Department of Electrical and Computer Engineering, Department of Internal Medicine, The University of Iowa, Iowa City, 52242, IA, USA.
⁷ 3312 Seamans Center for the Engineering Arts and Sciences, Iowa City, IA, 52242-1527, USA.
⁸ Center for Bioinformatics and Computational Biology, University of Iowa, 5017 Seamans Center, Iowa City, IA, 52242, USA.

PMID: 31208324
PMCID: PMC6580485
DOI: 10.1186/s12859-019-2929-8

Machine learning with the TCGA-HNSC dataset: improving usability by addressing inconsistency, sparsity, and high-dimensionality

Michael C Rendleman et al. BMC Bioinformatics. 2019.

. 2019 Jun 17;20(1):339.

doi: 10.1186/s12859-019-2929-8.

Authors

Michael C Rendleman¹, John M Buatti², Terry A Braun³, Brian J Smith⁴, Chibuzo Nwakama⁵, Reinhard R Beichel^{6

7}, Bart Brown⁸, Thomas L Casavant^{5

3}

Affiliations

¹ Department of Electrical and Computer Engineering, Center for Bioinformatics and Computational Biology, University of Iowa, 5017 Seamans Center, Iowa City, IA, 52242, USA. michael-rendleman@uiowa.edu.
² Department of Radiation Oncology, Carver College of Medicine, University of Iowa Carver College of Medicine, LL-W Pomerantz Family Pavilion, 200 Hawkins Drive, Iowa City, IA, 52242-1089, USA.
³ Department of Biomedical Engineering, Center for Bioinformatics and Computational Biology, University of Iowa, 5017 Seamans Center, Iowa City, IA, 52242, USA.
⁴ Department of Biostatistics, University of Iowa, 145 N. Riverside Drive, 100 CPHB, Iowa City, IA, 52242, USA.
⁵ Department of Electrical and Computer Engineering, Center for Bioinformatics and Computational Biology, University of Iowa, 5017 Seamans Center, Iowa City, IA, 52242, USA.
⁶ Iowa Institute for Biomedical Imaging, Department of Electrical and Computer Engineering, Department of Internal Medicine, The University of Iowa, Iowa City, 52242, IA, USA.
⁷ 3312 Seamans Center for the Engineering Arts and Sciences, Iowa City, IA, 52242-1527, USA.
⁸ Center for Bioinformatics and Computational Biology, University of Iowa, 5017 Seamans Center, Iowa City, IA, 52242, USA.

PMID: 31208324
PMCID: PMC6580485
DOI: 10.1186/s12859-019-2929-8

Abstract

Background: In the era of precision oncology and publicly available datasets, the amount of information available for each patient case has dramatically increased. From clinical variables and PET-CT radiomics measures to DNA-variant and RNA expression profiles, such a wide variety of data presents a multitude of challenges. Large clinical datasets are subject to sparsely and/or inconsistently populated fields. Corresponding sequencing profiles can suffer from the problem of high-dimensionality, where making useful inferences can be difficult without correspondingly large numbers of instances. In this paper we report a novel deployment of machine learning techniques to handle data sparsity and high dimensionality, while evaluating potential biomarkers in the form of unsupervised transformations of RNA data. We apply preprocessing, MICE imputation, and sparse principal component analysis (SPCA) to improve the usability of more than 500 patient cases from the TCGA-HNSC dataset for enhancing future oncological decision support for Head and Neck Squamous Cell Carcinoma (HNSCC).

Results: Imputation was shown to improve prognostic ability of sparse clinical treatment variables. SPCA transformation of RNA expression variables reduced runtime for RNA-based models, though changes to classifier performance were not significant. Gene ontology enrichment analysis of gene sets associated with individual sparse principal components (SPCs) are also reported, showing that both high- and low-importance SPCs were associated with cell death pathways, though the high-importance gene sets were found to be associated with a wider variety of cancer-related biological processes.

Conclusions: MICE imputation allowed us to impute missing values for clinically informative features, improving their overall importance for predicting two-year recurrence-free survival by incorporating variance from other clinical variables. Dimensionality reduction of RNA expression profiles via SPCA reduced both computation cost and model training/evaluation time without affecting classifier performance, allowing researchers to obtain experimental results much more quickly. SPCA simultaneously provided a convenient avenue for consideration of biological context via gene ontology enrichment analysis.

Keywords: Decision support; Dimensionality reduction; Gene ontology enrichment analysis; Machine learning; Unsupervised transformation; hnscc; tcga.

PubMed Disclaimer

Conflict of interest statement

The authors declare that they have no competing interests.

Figures

**Fig. 1**
Methods Flowchart. Flow diagram outlining the methods of this work. a Clinical data preprocessing and imputation. b RNA expression experiments and analysis

**Fig. 2**
Cumulative Percent Explained Variance. Percent explained variance from SPCA as it relates to the number of components retained. The dark vertical line indicates the value used for transforming RNA expression into the SPCA feature set for these experiments

**Fig. 3**
Importance Change with Imputation. Pre-imputation and post-imputation CIRF conditional variable importance for predicting two-year recurrence-free survival. Importance values are relative to the most important variable. Imputed treatment features are denoted with *, and several clinical variables are shown for comparison

**Fig. 4**
SPC Conditional Importance Values. Relative conditional variable importance values for the 10 SPCs, labeled X1–10. In cases where a very low importance is reported for an SPC, its effect on classifier performance is negligible

See this image and copyright information in PMC

References

1. Wang L, Jia Y, Jiang Z, Gao W, Wang B. FSCN1 is upregulated by SNAI2 and promotes epithelial to mesenchymal transition in head and neck squamous cell carcinoma. Cell Biol Int. 2017;41(8):833–841. doi: 10.1002/cbin.10786. - DOI - PubMed
1. Liu J, Wu Y, Wang Q, Liu X, Liao X, Pan J. Bioinformatic analysis of PFN2 dysregulation and its prognostic value in head and neck squamous carcinoma. (1744–8301 (Electronic)). 2018. - PubMed
1. Huang H, Lin C, Yang C, Ho C, Chang Y, Chang J, editors. An integrative analysis for Cancer studies. 2016 IEEE 16th international conference on bioinformatics and bioengineering (BIBE); 2016 31 Oct.-2 Nov. 2016.
1. Krempel R, Kulkarni P, Yim A, Lang U, Habermann B, Frommolt P. Integrative analysis and machine learning on cancer genomics data using the Cancer systems biology database (CancerSysDB) BMC Bioinformatics. 2018;19(1):156. doi: 10.1186/s12859-018-2157-7. - DOI - PMC - PubMed
1. Hu S, Yuan H, Li Z, Zhang J, Wu J, Chen Y, et al. Transcriptional response profiles of paired tumor-normal samples offer novel perspectives in pan-cancer analysis. Oncotarget. 2017;8(25):41334–41347. doi: 10.18632/oncotarget.17295. - DOI - PMC - PubMed

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Machine learning with the TCGA-HNSC dataset: improving usability by addressing inconsistency, sparsity, and high-dimensionality

Affiliations

Machine learning with the TCGA-HNSC dataset: improving usability by addressing inconsistency, sparsity, and high-dimensionality

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

References

MeSH terms

Substances

Grants and funding

LinkOut - more resources

Full Text Sources