Abstract / Summary
Abstract Background: Machine-learning methods can capture nonlinear associations in multidomain health data, but discrimination alone does not establish the reliability of probability-based disease stratification. Calibration, temporal transportability, threshold-dependent behavior, subgroup heterogeneity, and model interpretability require separate evaluation.
Objective: To develop and temporally validate an explainable machine-learning framework for stratifying prevalent physiciandiagnosed cardiovascular disease (CVD), while evaluating discrimination, probability calibration, decision-analytic behavior, subgroup robustness, and explanation stability.
Methods: National Health and Nutrition Examination Survey (NHANES) data from 2007–2018 were analyzed. Adults aged ≥ 20 years with a known composite CVD outcome were included. Prevalent CVD was defined from self-reported physician diagnoses of congestive heart failure, coronary heart disease, angina, heart attack, or stroke. NHANES 2007–2012 constituted the development period and 2013–2018 the chronologically later temporal-validation period. Logistic regression, random forest, XGBoost, and LightGBM were compared and tuned using development-only five-fold cross-validation. XGBoost was selected according to the predefined development ROC-AUC ranking criterion. A sigmoid calibration mapping and a Youden-derived analytical operating point were established using development data and locked before temporal evaluation. The uncalibrated and isotonic probability outputs were retained as sensitivity analyses. Temporal evaluation included ROC-AUC, average precision, Brier score, log loss, calibration intercept and slope, bootstrap confidence intervals, threshold-dependent metrics, decision-curve analysis, predictor-domain ablation, subgroup analysis, and SHAP-based explainability.
Results: The analytical cohort included 34,602 adults: 17,626 in development and 16,976 in temporal validation. The locked sigmoid-calibrated XGBoost achieved a temporal ROC-AUC of 0.847 (95% CI 0.839–0.856) and average precision of 0.403 (95% CI 0.382–0.427), compared with a prevalence baseline of approximately 0.113. The Brier score was 0.083 (95% CI 0.080–0.086), calibration intercept was −0.067 (95% CI −0.164 to 0.038), and calibration slope was 0.981 (95% CI 0.943–1.022). The uncalibrated XGBoost had a lower temporal Brier score (0.0809), indicating that the sigmoid layer did not improve overall probability accuracy. At the development-locked analytical operating point of 0.0760, sensitivity was 0.820, specificity 0.718, positive predictive value 0.270, and negative predictive value 0.969. Operating characteristics varied substantially by age. Age was the dominant SHAP predictor, and original-predictor importance rankings were stable across temporal cycles (pairwise Spearman ρ= 0.982–0.996).