Abstract / Summary
Abstract Background Infections are a major cause of mortality in hemodialysis patients, but early risk stratification remains difficult. This study developed and validated an interpretable ML model using routine baseline laboratory data to predict short-term infection risk. Methods A retrospective cohort of 622 hemodialysis patients with uremia was analyzed. The primary outcome was a confirmed infectious event (infection, sepsis, or pneumonia) occurring within a 2-month follow-up. To rigorously prevent data leakage, the cohort was strictly partitioned into training (70%) and test (30%) sets prior to missing value imputation and any downstream preprocessing. A LightGBM-based forward feature selection was executed exclusively within the training set. Nine ML algorithms were evaluated, and model robustness was validated via nested cross-validation, repeated cross-validation, and bootstrap optimism correction. Interpretability was achieved using SHapley Additive exPlanations (SHAP). Results Subsequent infections occurred in 157 patients (25.2%). A parsimonious six-biomarker signature (CysC, HCT, TP, MONO_pct, MYO, ADA) was identified. The K-Nearest Neighbors (KNN) and LightGBM models demonstrated robust discriminative performance (Test AUROC: 0.727 and 0.719, respectively). Nested cross-validation confirmed generalization stability (e.g., ANN mean AUC: 0.722). SHAP analysis revealed elevated Cystatin C and lowered Hematocrit as the strongest contributors to the model’s risk ranking. Although the model showed high ranking ability, the default 0.5 threshold yielded low sensitivity; a prespecified F2‑optimized threshold of 0.30 improved sensitivity to 0.723 in the test set. Conclusion This interpretable machine learning‑based predictive model, using routinely available laboratory data, may serve as a potential adjunctive decision‑support tool for stratifying short‑term infection risk. External validation is required before clinical implementation.