Impact of censoring on learning Bayesian networks in survival modelling
- PMID: 19833488
- DOI: 10.1016/j.artmed.2009.08.001
Impact of censoring on learning Bayesian networks in survival modelling
Abstract
Objective: Bayesian networks are commonly used for presenting uncertainty and covariate interactions in an easily interpretable way. Because of their efficient inference and ability to represent causal relationships, they are an excellent choice for medical decision support systems in diagnosis, treatment, and prognosis. Although good procedures for learning Bayesian networks from data have been defined, their performance in learning from censored survival data has not been widely studied. In this paper, we explore how to use these procedures to learn about possible interactions between prognostic factors and their influence on the variate of interest. We study how censoring affects the probability of learning correct Bayesian network structures. Additionally, we analyse the potential usefulness of the learnt models for predicting the time-independent probability of an event of interest.
Methods and materials: We analysed the influence of censoring with a simulation on synthetic data sampled from randomly generated Bayesian networks. We used two well-known methods for learning Bayesian networks from data: a constraint-based method and a score-based method. We compared the performance of each method under different levels of censoring to those of the naive Bayes classifier and the proportional hazards model. We did additional experiments on several datasets from real-world medical domains. The machine-learning methods treated censored cases in the data as event-free.
Results: We report and compare results for several commonly used model evaluation metrics. On average, the proportional hazards method outperformed other methods in most censoring setups. As part of the simulation study, we also analysed structural similarities of the learnt networks. Heavy censoring, as opposed to no censoring, produces up to a 5% surplus and up to 10% missing total arcs. It also produces up to 50% missing arcs that should originally be connected to the variate of interest.
Conclusion: Presented methods for learning Bayesian networks from data can be used to learn from censored survival data in the presence of light censoring (up to 20%) by treating censored cases as event-free. Given intermediate or heavy censoring, the learnt models become tuned to the majority class and would thus require a different approach.
Similar articles
-
Learning Bayesian networks from survival data using weighting censored instances.J Biomed Inform. 2010 Aug;43(4):613-22. doi: 10.1016/j.jbi.2010.03.005. Epub 2010 Mar 21. J Biomed Inform. 2010. PMID: 20332035
-
Bayesian networks for multivariate data analysis and prognostic modelling in cardiac surgery.Stud Health Technol Inform. 2007;129(Pt 1):596-600. Stud Health Technol Inform. 2007. PMID: 17911786
-
Prognostic Bayesian networks I: rationale, learning procedure, and clinical use.J Biomed Inform. 2007 Dec;40(6):609-18. doi: 10.1016/j.jbi.2007.07.003. Epub 2007 Jul 25. J Biomed Inform. 2007. PMID: 17704008
-
Bottom-up learning of explicit knowledge using a Bayesian algorithm and a new Hebbian learning rule.Neural Netw. 2011 Apr;24(3):219-32. doi: 10.1016/j.neunet.2010.12.002. Epub 2010 Dec 16. Neural Netw. 2011. PMID: 21239141 Review.
-
Pseudo-observations in survival analysis.Stat Methods Med Res. 2010 Feb;19(1):71-99. doi: 10.1177/0962280209105020. Epub 2009 Aug 4. Stat Methods Med Res. 2010. PMID: 19654170 Review.
Cited by
-
A probabilistic analysis of completely excised high-grade soft tissue sarcomas of the extremity: an application of a Bayesian belief network.Ann Surg Oncol. 2012 Sep;19(9):2992-3001. doi: 10.1245/s10434-012-2345-z. Epub 2012 Apr 20. Ann Surg Oncol. 2012. PMID: 22526900 Free PMC article.
-
Learning rule sets from survival data.BMC Bioinformatics. 2017 May 30;18(1):285. doi: 10.1186/s12859-017-1693-x. BMC Bioinformatics. 2017. PMID: 28558674 Free PMC article.
-
A Naive Bayes machine learning approach to risk prediction using censored, time-to-event data.Stat Med. 2015 Sep 20;34(21):2941-57. doi: 10.1002/sim.6526. Epub 2015 May 18. Stat Med. 2015. PMID: 25980520 Free PMC article.
-
A novel dynamic Bayesian network approach for data mining and survival data analysis.BMC Med Inform Decis Mak. 2022 Sep 22;22(1):251. doi: 10.1186/s12911-022-02000-7. BMC Med Inform Decis Mak. 2022. PMID: 36138394 Free PMC article.
-
Application of machine learning algorithms for clinical predictive modeling: a data-mining approach in SCT.Bone Marrow Transplant. 2014 Mar;49(3):332-7. doi: 10.1038/bmt.2013.146. Epub 2013 Oct 7. Bone Marrow Transplant. 2014. PMID: 24096823 Review.
Publication types
MeSH terms
LinkOut - more resources
Full Text Sources
Research Materials