Abstract / Summary
Background: Adjuvant therapy decision in early-stage node-negative breast cancer remains one of the most challenging problems in oncology. Although most patients in this population achieve favorable outcomes with endocrine therapy alone, a subset remains at risk of relapse and may benefit from chemotherapy. Accurate identification of high-risk and low-risk patients is therefore essential to prevent both undertreatment and overtreatment. Methods: In this study, machine learning models were assessed using MSK-IMPACT (n=645) and METABRIC (n=503) cohorts comprising node-negative stage I-II breast cancer patients. Three experimental settings were evaluated: training on MSK-IMPACT and testing on METABRIC, the reverse direction, and five-fold cross-validation on the combined cohort. Sensitivity analysis and data harmonization were applied to identify and mitigate cohort-specific bias. Survival analysis, SHapley Additive exPlanations (SHAP) interpretability, and ablation analyses were performed to validate clinical relevance and characterize feature contributions. Results: The combined cohort model achieved an area under the receiver operating characteristic curve (AUC) of 0.788{+/-}0.028 across five-fold cross-validation, while bidirectional cross-cohort validation yielded AUC of 0.741 when testing on MSK-IMPACT and AUC of 0.714 when testing on METABRIC. Survival validation on METABRIC out-of-fold predictions from the combined cohort model showed hazard ratios of 3.01 (95\% CI: 2.28-3.98) for disease-free survival and 2.35 (95\% CI: 1.79-3.09) for overall survival, consistent across Luminal A, Luminal B, and Triple-Negative subtypes. Conclusion: Validated across two geographically independent cohorts, the proposed model enables cost-effective, gene expression-free identification of patients at high-risk of distant relapse and may facilitate adjuvant chemotherapy decision-making in early-stage node-negative breast cancer.