Forecasting incidence of infectious diarrhea using random forest in Jiangsu Province, China
- PMID: 32171261
- PMCID: PMC7071679
- DOI: 10.1186/s12879-020-4930-2
Forecasting incidence of infectious diarrhea using random forest in Jiangsu Province, China
Abstract
Background: Infectious diarrhea can lead to a considerable global disease burden. Thus, the accurate prediction of an infectious diarrhea epidemic is crucial for public health authorities. This study was aimed at developing an optimal random forest (RF) model, considering meteorological factors used to predict an incidence of infectious diarrhea in Jiangsu Province, China.
Methods: An RF model was developed and compared with classical autoregressive integrated moving average (ARIMA)/X models. Morbidity and meteorological data from 2012 to 2016 were used to construct the models and the data from 2017 were used for testing.
Results: The RF model considered atmospheric pressure, precipitation, relative humidity, and their lagged terms, as well as 1-4 week lag morbidity and time variable as the predictors. Meanwhile, a univariate model ARIMA (1,0,1)(1,0,0)52 (AIC = - 575.92, BIC = - 558.14) and a multivariable model ARIMAX (1,0,1)(1,0,0)52 with 0-1 week lag precipitation (AIC = - 578.58, BIC = - 578.13) were developed as benchmarks. The RF model outperformed the ARIMA/X models with a mean absolute percentage error (MAPE) of approximately 20%. The performance of the ARIMAX model was comparable to that of the ARIMA model with a MAPE reaching approximately 30%.
Conclusions: The RF model fitted the dynamic nature of an infectious diarrhea epidemic well and delivered an ideal prediction accuracy. It comprehensively combined the synchronous and lagged effects of meteorological factors; it also integrated the autocorrelation and seasonality of the morbidity. The RF model can be used to predict the epidemic level and has a high potential for practical implementation.
Keywords: Forecasting; Infectious diarrhea; Random forest.
Conflict of interest statement
The authors declare that they have no competing interests.
Figures
References
-
- GBD 2015 Disease and Injury Incidence and Prevalence Collaborators Global, regional, and national incidence, prevalence, and years lived with disability for 310 diseases and injuries, 1990–2015: a systematic analysis for the Global Burden of Disease Study 2015. Lancet. 2016;388(10053):1545–1602. - PMC - PubMed
-
- GBD 2015 Mortality and Causes of Death Collaborators Global, regional, and national life expectancy, all-cause mortality, and cause-specific mortality for 249 causes of death, 1980–2015: a systematic analysis for the Global Burden of Disease Study 2015. Lancet. 2016;388(10053):1459–1544. - PMC - PubMed
-
- Zhang P, Zhang J. Surveillance on other infectious diarrheal diseases in China from 2014 to 2015. Chin J Epidemiol 2017;38(4):424–430.(in Chinese). - PubMed
-
- Zhang Y, Bi P, Hiller JE, Sun Y, Ryan P. Climate variations and bacillary dysentery in northern and southern cities of China. J Inf Secur. 2007;55(2):194–200. - PubMed
MeSH terms
Grants and funding
- BE2017749/National Major Science and Technology Projects of China
- BK20151595/Jiangsu Provincial natural science foundation
- QNRC2016542/Jiangsu Provincial Medical Youth Talent
- ZDXK A2016008/Key Medical Discipline of Epidemiology
- 2018ZX10714-002/National Science and Technology Infrastructure Program (CN)
LinkOut - more resources
Full Text Sources
Medical
