Abstract / Summary
Background Postoperative pulmonary complications (PPCs) significantly impair recovery after lung cancer surgery. We developed and validated an ensemble machine learning (ML) model to identify high-risk patients across the perioperative period. Methods We retrospectively analyzed data from 3,476 patients undergoing anatomic lung resection between 2016 and 2020 (Cohort 1). Predictors were selected through clinical expertise, statistical significance, and the Boruta algorithm. Seven ML models were developed, with the soft Voting Classifier (integrating Random Forest, CatBoost, and XGBoost) identified as the optimal approach. A prospective temporal validation cohort (n = 171) was also evaluated. Performance was assessed by AUROC with patient-level bootstrap 95% CIs, Brier score, calibration, and decision-curve analysis.. Results The overall incidence of PPCs was 15.9%. In internal validation, the Voting Classifier ranked highest among eight models for overall PPCs (AUROC 0.7216; 95% CI 0.7044–0.7369), comparable to XGBoost and CatBoost with overlapping confidence intervals, and numerically higher than logistic regression (0.7009). In temporal validation, the model maintained moderate discriminative ability (AUROC 0.6366; 95% CI 0.5179–0.7551). Despite a surgical paradigm shift (VATS from 73.8% to 98.2%), the model retained discriminative ability for postoperative pneumonia (AUROC 0.7434; 95% CI 0.5887–0.8759), the strongest of the three subtype predictions (internal AUROC 0.7799); discrimination was inconclusive for prolonged air leak (0.6790; 0.4673–0.8796) . Conclusions The Voting Classifier provides an interpretable tool for predicting PPCs and guiding perioperative risk stratification, enabling targeted interventions for high-risk patients. The model was developed in patients undergoing upfront anatomic lung resection and does not apply to those receiving neoadjuvant therapy Ethical Registration ChiCTR2500102082 on May 8 2025