Despite growing enthusiasm for artificial intelligence (AI) in epidemiology, its added value for short-term forecasting remains uncertain. In this method, using CI5plus international cancer registry data (1990–2017) enriched with World Bank urban environment indicators, we benchmark one-year-ahead cancer incidence forecasting under strict temporal validation (train 1990–2012; validation 2013–2015; test 2016–2017). We compare naive temporal base lines (last observation, three-year moving average), classical time-series models (ARIMA, exponential smoothing), machine learning (ridge regression, random forest), and a deep learning model (multilayer perceptron). Model differences are assessed using Diebold–Mariano tests and bootstrap confidence intervals, and robustness is verified through rolling-origin evaluation. Regarding the results, across country–sex–site strata, the last-observation baseline consistently achieved the best test performance, significantly outperforming all competing models (Diebold–Mariano p < 0.05 for all pairwise comparisons on MAE). These results were robust across three rolling-origin windows. Urban environment covariates provided negligible incremental predictive value beyond recent incidence history. In conclusion, for short-horizon cancer incidence forecasting with highly persistent series, strong naive baselines are difficult to beat. Rigorous temporal evaluation, statistical comparison, and appropriate baselines are essential for credible claims of AI benefit in epidemiological forecasting.
Copyrights © 2026