Comparative Study

. 2010 Jan 7:11:15.

doi: 10.1186/1471-2164-11-15.

Comparative analysis of missing value imputation methods to improve clustering and interpretation of microarray experiments

Magalie Celton¹, Alain Malpertuy, Gaëlle Lelandais, Alexandre G de Brevern

Affiliations

PMID: 20056002
PMCID: PMC2827407
DOI: 10.1186/1471-2164-11-15

Comparative Study

Comparative analysis of missing value imputation methods to improve clustering and interpretation of microarray experiments

Magalie Celton et al. BMC Genomics. 2010.

. 2010 Jan 7:11:15.

doi: 10.1186/1471-2164-11-15.

Authors

Magalie Celton¹, Alain Malpertuy, Gaëlle Lelandais, Alexandre G de Brevern

Affiliation

¹ INSERM UMR-S 726, Equipe de Bioinformatique Génomique et Moléculaire, DSIMB, Université Paris Diderot-Paris 7, 2 place Jussieu, Paris, France.

PMID: 20056002
PMCID: PMC2827407
DOI: 10.1186/1471-2164-11-15

Abstract

Background: Microarray technologies produced large amount of data. In a previous study, we have shown the interest of k-Nearest Neighbour approach for restoring the missing gene expression values, and its positive impact of the gene clustering by hierarchical algorithm. Since, numerous replacement methods have been proposed to impute missing values (MVs) for microarray data. In this study, we have evaluated twelve different usable methods, and their influence on the quality of gene clustering. Interestingly we have used several datasets, both kinetic and non kinetic experiments from yeast and human.

Results: We underline the excellent efficiency of approaches proposed and implemented by Bo and co-workers and especially one based on expected maximization (EM_array). These improvements have been observed also on the imputation of extreme values, the most difficult predictable values. We showed that the imputed MVs have still important effects on the stability of the gene clusters. The improvement on the clustering obtained by hierarchical clustering remains limited and, not sufficient to restore completely the correct gene associations. However, a common tendency can be found between the quality of the imputation method and the gene cluster stability. Even if the comparison between clustering algorithms is a complex task, we observed that k-means approach is more efficient to conserve gene associations.

Conclusions: More than 6.000.000 independent simulations have assessed the quality of 12 imputation methods on five very different biological datasets. Important improvements have so been done since our last study. The EM_array approach constitutes one efficient method for restoring the missing expression gene values, with a lower estimation error level. Nonetheless, the presence of MVs even at a low rate is a major factor of gene cluster instability. Our study highlights the need for a systematic assessment of imputation methods and so of dedicated benchmarks. A noticeable point is the specific influence of some biological dataset.

PubMed Disclaimer

Figures

**Figure 1**
**Principle of the method**. The initial data matrix is analyzed. Each gene associated to at least one missing value (*in pink*) is excluded given a *Reference matrix* without any missing value. Then missing values are simulated (*in red*) with a fixed rate τ. This rate τ goes from 0.5% to 50% of missing values by step of 0.5%. 100 independent simulations are done each time. Missing values are then imputed (*in blue*) for each simulations by the selected methods. *RMSE* is computed between the estimated values of missing values and their true values.

**Figure 2**
**Example of three methods**. Distribution of predicted values (y-axis) in regards to true values (x-axis). Estimation of the missing values has been done (a) by *kNN* approach, (b) *EM_gene* and (c) *EM_array*. The dataset used is the Bohen set with τ values ranging from 0.5% to 50% of missing values with a step of 0.5. 10 independent simulations have been done for each τ value.

**Figure 3**
**Missing value imputation**. RMSE value for (a) G_Heatsubset and (b) for OS for rate of missing value going from 0.5% to 50% by step of 0.5%. (b) 100 independent simulations are done at each level.

**Figure 4**
**Extreme values (representing 1% of the missing values)**. Evolution of RMSE according to τ ranging (a) from 0.5% to 30% of the extreme values for the Bohen dataset and (b) from 0.5% to 50% of the extreme values) for the Ogawa dataset.

**Figure 5**
*CPP* of hierarchical clustering approach algorithm. (a) with complete, average, ward and McQuitty algorithm for OS with *kNN* and (b) with Ward algorithm for Ogawa dataset for the different imputation methods.

**Figure 6**
**Summary of the comparison**.

See this image and copyright information in PMC

Cited by

Trends in IT Innovation to Build a Next Generation Bioinformatics Solution to Manage and Analyse Biological Big Data Produced by NGS Technologies.
de Brevern AG, Meyniel JP, Fairhead C, Neuvéglise C, Malpertuy A. de Brevern AG, et al. Biomed Res Int. 2015;2015:904541. doi: 10.1155/2015/904541. Epub 2015 Jun 1. Biomed Res Int. 2015. PMID: 26125026 Free PMC article. Review.
Classifying Incomplete Gene-Expression Data: Ensemble Learning with Non-Pre-Imputation Feature Filtering and Best-First Search Technique.
Yan Y, Dai T, Yang M, Du X, Zhang Y, Zhang Y. Yan Y, et al. Int J Mol Sci. 2018 Oct 30;19(11):3398. doi: 10.3390/ijms19113398. Int J Mol Sci. 2018. PMID: 30380746 Free PMC article.
Latent triple trajectories of substance use as predictors for the onset of antisocial personality disorder among urban African American and Puerto Rican adults: A 22-year longitudinal study.
Lee JY, Pahl K, Kim W. Lee JY, et al. Subst Abus. 2022;43(1):442-450. doi: 10.1080/08897077.2021.1946890. Subst Abus. 2022. PMID: 35073243 Free PMC article.
Shrinkage regression-based methods for microarray missing value imputation.
Wang H, Chiu CC, Wu YC, Wu WS. Wang H, et al. BMC Syst Biol. 2013;7 Suppl 6(Suppl 6):S11. doi: 10.1186/1752-0509-7-S6-S11. Epub 2013 Dec 13. BMC Syst Biol. 2013. PMID: 24565159 Free PMC article.
Gene set enrichment analysis: performance evaluation and usage guidelines.
Hung JH, Yang TH, Hu Z, Weng Z, DeLisi C. Hung JH, et al. Brief Bioinform. 2012 May;13(3):281-91. doi: 10.1093/bib/bbr049. Epub 2011 Sep 7. Brief Bioinform. 2012. PMID: 21900207 Free PMC article. Review.

See all "Cited by" articles

References

1. Liolios K, Tavernarakis N, Hugenholtz P, Kyrpides NC. The Genomes On Line Database (GOLD) v.2: a monitor of genome projects worldwide. Nucleic Acids Res. 2006. pp. D332–334. - DOI - PMC - PubMed
1. Bernal A, Ear U, Kyrpides N. Genomes OnLine Database (GOLD): a monitor of genome projects world-wide. Nucleic Acids Res. 2001;29(1):126–127. doi: 10.1093/nar/29.1.126. - DOI - PMC - PubMed
1. Muro S, Takemasa I, Oba S, Matoba R, Ueno N, Maruyama C, Yamashita R, Sekimoto M, Yamamoto H, Nakamori S. et al.Identification of expressed genes linked to malignancy of human colorectal carcinoma by parametric clustering of quantitative expression data. Genome Biol. 2003;4(3):R21. doi: 10.1186/gb-2003-4-3-r21. - DOI - PMC - PubMed
1. Perou CM, Sorlie T, Eisen MB, Rijn M van de, Jeffrey SS, Rees CA, Pollack JR, Ross DT, Johnsen H, Akslen LA. et al.Molecular portraits of human breast tumours. Nature. 2000;406(6797):747–752. doi: 10.1038/35021093. - DOI - PubMed
1. Statnikov A, Aliferis CF, Tsamardinos I, Hardin D, Levy S. A comprehensive evaluation of multicategory classification methods for microarray gene expression cancer diagnosis. Bioinformatics. 2005;21(5):631–643. doi: 10.1093/bioinformatics/bti033. - DOI - PubMed

Publication types

Actions
Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
Molecular Biology Databases
- Saccharomyces Genome Database

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Comparative analysis of missing value imputation methods to improve clustering and interpretation of microarray experiments

Affiliation

Comparative analysis of missing value imputation methods to improve clustering and interpretation of microarray experiments

Authors

Affiliation

Abstract

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

LinkOut - more resources

Full Text Sources

Molecular Biology Databases

Abstract

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

Related information

LinkOut - more resources

Full Text Sources

Molecular Biology Databases