Abstract / Summary
Background: Preoperative nodal assessment in colorectal cancer (CRC) remains imperfect, and uncertainty regarding pathological lymph node metastasis (LNM) may persist despite conventional imaging. Machine learning using routinely available preoperative information may provide complementary patient-level risk estimates without requiring radiomic feature extraction or molecular testing.
Objective: This study aimed to develop and externally validate machine learning models for preoperative estimation of pathologically confirmed LNM risk in patients with newly diagnosed CRC undergoing direct radical surgery without preoperative anticancer treatment, and to assess whether multivariable prediction adds information beyond clinical tumor stage (cT) alone.
Methods: This multicenter retrospective study included 2725 patients from 2 hospitals. The internal cohort (n=1074) was split into training (n=753) and internal test (n=321) sets, while 1651 patients from a second center formed the external validation cohort. Candidate predictors were routinely available preoperative clinical, laboratory, tumor-marker, imaging-based, and endoscopic biopsy variables. Predictors were selected using univariable and multivariable logistic regression (LR). Seven algorithms were evaluated: random forest (RF), support vector machine (SVM), LR, decision tree (DT), extreme gradient boosting (XGBoost), naive Bayes (NB), and light gradient boosting machine (LightGBM). Performance was assessed using discrimination, threshold-dependent classification metrics, calibration, and decision curve analysis (DCA). Shapley additive explanations (SHAP) were used for interpretation. Sensitivity analyses compared the selected predictors with all candidate predictors and cT alone.
Results: Six predictors were retained: BMI, preoperative carcinoembryonic antigen, primary tumor site, cT, histological type, and tumor differentiation. RF was selected based on internal test performance. The area under the receiver operating characteristic curve (AUROC) was 0.803 (95% CI 0.756-0.849) in the internal test set and 0.776 (95% CI 0.755-0.797) in external validation. In external validation, accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were 0.687, 0.677, 0.693, 0.569, and 0.781, respectively. The Brier score was 0.190, and the calibration slope was 1.061. cT contributed most in SHAP analysis; however, the 6-predictor RF model showed significantly higher discrimination than cT alone in both the internal test set (AUROC 0.803 vs 0.712; ΔAUROC=0.091; P<.001) and external validation cohort (AUROC 0.776 vs 0.707; ΔAUROC=0.069; P<.001). Discrimination did not differ significantly between the 6-predictor and all-predictor RF models.