Abstract / Summary
The COVID-19 pandemic accelerated the development of artificial intelligence (AI), machine learning (ML), and deep learning (DL) for clinical decision support, public-health surveillance, and pandemic response. This systematic review evaluates AI-based studies addressing COVID-19 diagnosis, prognosis, outbreak forecasting, remote monitoring, digital epidemic intelligence, and multimodal applications, emphasizing dataset characteristics, validation maturity, generalizability, fairness, privacy, and deployment readiness. Following PRISMA 2020, studies published between December 2019 and 31 January 2024 were retrieved from IEEE Xplore, ACM Digital Library, SpringerLink, ScienceDirect, Wiley Online Library, and arXiv using database-adapted search strategies. Of 110 identified records, 57 met the eligibility criteria and underwent a six-domain CASP-informed methodological appraisal and validation-maturity assessment. Quantitative synthesis characterized application domain, publication timing, dataset geographic provenance and scale, principal model family, imaging modality, and validation strategy. Diagnostic imaging and clinical diagnosis constituted the largest evidence domain, followed by prognosis, forecasting, remote monitoring/IoT, OSN/NLP surveillance, and multimodal or hybrid systems. Most studies remained at the development-validation stage: 46 of 57 (80.7%) relied primarily on internal splitting or cross-validation, whereas only 11 (19.3%) reached external, cross-dataset, or multicenter validation. Although many models reported strong internal performance, especially for chest X-ray and CT-based diagnosis and clinical-risk prediction, generalizability was constrained by retrospective or geographically restricted datasets, source heterogeneity, leakage, shortcut learning, hidden confounding, temporal drift, and subgroup bias. OSN- and IoT-based approaches offered scalable surveillance and monitoring but remained limited by noisy signals, demographic bias, privacy risks, and implementation variability. Overall, the evidence reveals a persistent gap between technical performance and real-world readiness. Future pandemic AI should prioritize independent and multicenter validation, transparent dataset provenance, fairness-aware and explainable evaluation, multimodal integration, calibration under distribution shift, and privacy-preserving approaches such as federated learning.