Abstract / Summary
Abstract Background Immune checkpoint inhibitor-associated myocarditis is rare but potentially fatal. We developed and internally evaluated machine-learning models for camrelizumab-associated myocarditis, focusing on rare-event uncertainty, predictor timing, and calibration. Methods We retrospectively analyzed 2,774 patients who first received camrelizumab between March 2021 and May 2023. The outcome was myocarditis defined by the 2022 European Society of Cardiology cardio-oncology criteria. Laboratory predictors were the first available results recorded before the first camrelizumab administration (index date); no fixed look-back window was applied. Troponin, B-type natriuretic peptide (BNP)/N-terminal pro-BNP, and vital signs were excluded from the primary predictor set. A random forest (RF) was the parsimonious reference model; regularized logistic regression and exploratory probability fusion with extreme gradient boosting (XGBoost) were comparators. Models were evaluated using repeated stratified five-fold cross-validation (three repeats) with fixed settings. Overall estimates averaged three cross-fitted probabilities per patient, and 95% confidence intervals (CIs) were obtained from 2,000 patient-level bootstrap resamples. Results Forty-nine patients (1.77%) met the outcome. Across 15 folds, RF mean (standard deviation) area under the receiver operating characteristic curve (ROC-AUC) was 0.8501 (0.0501), area under the precision-recall curve (PR-AUC) 0.3552 (0.1429), and Brier score 0.0151 (0.0012). Patient-averaged estimates were ROC-AUC 0.8524 (95% CI, 0.7938–0.9015), PR-AUC 0.2937 (95% CI, 0.1708–0.4477), and Brier score 0.0150 (95% CI, 0.0112–0.0195). Across three fusion weights, ROC-AUC and PR-AUC were slightly higher and Brier scores slightly lower than for RF. For the 70:30 fusion, all paired 95% CIs for the ROC-AUC, PR-AUC, and Brier-score differences included zero. In an exploratory temporal evaluation of patients with index dates from 1 January to 9 May 2023 (310 patients; six events), ROC-AUC was 0.8410 (95% CI, 0.6022–0.9902), but sensitivity was 0.1667 at the training-derived threshold. Conclusions RF showed promising internal discrimination. The consistent numerical gains with fusion support further evaluation, but neither model is ready for clinical use because only 49 events were available and no external cohort was evaluated.