Abstract / Summary
Objective: To determine how predictive performance varies across artificial intelligence model architectures for predicting unplanned readmission and mortality in the intensive care unit (ICU), evaluated either as separate or composite outcomes. Methods: A systematic review was conducted following PRISMA guidelines. Searches were performed in Scopus, Web of Science, PubMed, and IEEE Xplore for studies published between 2020 and 2025. Studies reporting predictive performance for both outcomes and providing a clearly described validation strategy were included. Studies involving COVID-19 patients were excluded. Risk of bias was assessed using PROBAST, and meta-analysis was performed using a random-effects model. Results: From 4105 identified records, 8 studies met the inclusion criteria. Because pooling AUC-ROC across studies confounds architectural differences with differences between cohorts, outcome definitions, and validation designs, this review prioritizes three paired within-study comparisons, each free of that confounding. First, when each study’s own proposed architecture was compared with that same study’s logistic-regression baseline on the identical cohort and outcome, the pooled advantage was +0.027 AUC-ROC (95% CI: +0.010 to +0.044; k=4; I2 = 0.0%)—a small, consistent gain, far more homogeneous than any cross-study pooled estimate. Second, within a single study evaluating eight identical architectures on two databases (MIMIC-III and eICU-CRD), changing the database moved AUC-ROC by 0.068 on average (range 0.041–0.101)—2.5 times the pooled architecture effect—while the ranking of architectures was preserved. Third, in the three studies reporting both an internal and an external validation cohort, external validation reduced AUC-ROC by 0.055–0.073 relative to internal validation. Against this background, cross-study pooled estimates are reported as exploratory context only. AUC-ROC standard errors were reconstructed at the estimate level wherever possible (reported confidence interval, or the Hanley–McNeil sampling-variance formula, corrected for each study’s actual evaluation-set size where a train/test split was used) rather than assumed to be uniform; one study (Sun M. et al., which does not report event counts and shows independent evidence of an atypically high evaluation-cohort prevalence) was excluded from pooling. Pooled on the logit scale to keep confidence intervals within the mathematically admissible [0,1] range, the primary analysis (one prespecified representative estimate per study, k=7) yielded AUC-ROC = 0.814 (95% CI: 0.746–0.868; I2 = 98.5%), and an exploratory analysis of the remaining 68 model/outcome estimates yielded a similar point estimate (0.807, 95% CI: 0.786–0.826; I2 = 98.8%). This heterogeneity is not an artifact of understated standard errors: inflating every Hanley–McNeil standard error by a factor of five reduced I2 only to 71–75%. Three studies were judged to be at low risk of bias, five raised some concerns, and none was at high risk. Limitations: The small number of included studies (k=7 in the primary pooled analysis) and substantial methodological and outcome-definition heterogeneity limit comparability. Four of the seven pooled studies share Amsterdam UMC as their underlying data source and two share MIMIC, so even this one-estimate-per-study analysis is not fully independent; restricting it to one study per database family shifts the pooled estimate to 0.857 (k=5), underscoring that it should be read as exploratory context rather than an effect estimate. Only 17 of 78 AUC-ROC estimates had a reported confidence interval; standard errors for the remaining estimates were reconstructed from sample and event counts where available (51/78). This review was not prospectively registered in PROSPERO. Interpretation: The most defensible quantitative finding of this review is that, within individual studies, the choice of source database moves AUC-ROC more than the choice of model architecture. Deep learning architectures achieved the highest cross-study point estimates, but the paired comparisons show a far smaller architectural advantage once cohort differences are held constant, and the evidence remains predominantly based on internal validation; pooled AUC-ROC figures should not be read as a stable estimate of expected performance. Prospective multicenter studies with robust external validation, explicit reporting of evaluation-set size, and harmonized outcome definitions are needed to establish clinical utility.