Abstract / Summary
Objectives: Machine learning models for graft-versus-host disease (GVHD) after allogeneic hematopoietic stem cell transplantation (allo-HSCT) are rarely evaluated externally or for calibration. We developed explainable models for grade II-IV acute and moderate-to-severe chronic GVHD and evaluated their transportability across multiple cohorts. Materials and methods: Gradient boosting models were trained on three CIBMTR registry-derived datasets (N = 2,509) and externally evaluated in six cohorts (N = 14,788), including a Middle Eastern cohort. We assessed discrimination (AUROC), calibration (logistic recalibration intercept and slope; Brier score), and SHAP feature explanations. Results: Internal discrimination was higher for chronic (AUROC 0.736) than acute GVHD (0.624); chronic calibration was good (intercept +0.258, slope 1.073), whereas the acute model was underdispersed (slope 1.802). Externally, the chronic model retained modest, above-chance discrimination (AUROC 0.54-0.64), whereas acute performance fell toward chance (0.52-0.57). Calibration drifted unpredictably, overestimating risk in 9 of 10 evaluable cohort-outcome combinations (intercepts -0.78 to +0.15; slopes 0.12 to 1.98; Brier 0.23-0.27). Discussion: Routine registry variables carry a transportable signal for chronic GVHD, but acute GVHD likely requires additional biological or clinical predictors. Neither the direction nor the magnitude of calibration drift could be anticipated from training performance. Conclusion: Internal performance does not guarantee external utility. Cohort-specific recalibration and routine calibration reporting are required before any clinical use; this study offers a reusable template for multi-cohort external evaluation of clinical prediction models.