Abstract / Summary
ABSTRACT Background Tuberculosis (TB) remains a significant global health problem, particularly in low‐ and middle‐income countries (LMICs), where accurate prediction of TB incidence is crucial for effective public health interventions. Objective This review assessed and compared statistical, machine learning (ML), and hybrid models for forecasting TB incidence. We also propose a framework to guide model selection based on resources and data. Methods A systematic review followed PRISMA 2020 guidelines with protocol in PROSPERO (CRD42025633162). We searched five databases for studies from 2013 to 2024. Studies using TB incidence forecasting models with quantitative metrics were eligible. Due to heterogeneity ( I 2 = 99.89%), we conducted narrative synthesis. Subgroup analyses examined geographic variation, with risk of bias assessed using ROBIS tool. Publication bias was evaluated through funnel plots and Egger's test. Results We included 29 studies. Hybrid models, particularly ARIMA‐LSTM, demonstrated superior forecasting accuracy in most contexts, with reported mean absolute percentage error (MAPE) values ranging from 4.06% to 11.2%. In contrast, standalone ML models showed MAPE between 8.88% and 26.93%, and statistical models ranged from 3.77% to 29.37%. However, only 36% of studies validated models using external datasets, and fewer than 30% incorporated exogenous variables. Subgroup analysis revealed geographic disparity: Within China, single models slightly outperformed hybrids (mean MAPE 8.76% vs. 12.95%), whereas outside China hybrids were markedly better (5.42% vs. 28.24%). Publication bias was suggested by funnel plot asymmetry (Egger's test p = 0.22). Risk of bias was moderate to high in most studies, primarily due to lack of external validation. Conclusion Hybrid models, especially ARIMA‐LSTM, often provide accurate TB incidence forecasts, but performance is highly context‐dependent. Future research must prioritize external validation, integration of exogenous variables, and model interpretability. The proposed evidence‐based framework can assist model selection in resource‐limited LMIC settings.