Abstract / Summary
Abstract Background: Bone metastasis is a major complication in breast cancer with substantial impact on patient prognosis. Machine learning applied to transcriptomic data has been proposed as a promising approach for risk stratification, but concerns remain regarding reproducibility, data leakage, and generalizability across independent test sets. Methods: We systematically evaluated four machine learning approaches: Random Forest (RF), Elastic Net (GLMNET), XGBoost, and L2-regularized logistic regression (LR) for predicting bone metastasis from microarray gene expression data. We used the GSE2034 cohort (n = 286, 69 bone metastasis events), strictly separated training (70%, 49 events) and test (30%, 20 events) sets, and applied all feature selection, class weighting, and model fitting exclusively within the training set. Model performance was assessed using five-fold cross-validated out-of-fold (CV-OOF) area under the receiver operating characteristic curve (AUC), independent test-set AUC with bootstrap 95% confidence intervals, precision-recall AUC (PR-AUC), Brier score, isotonic probability calibration, decision curve analysis (DCA), and SHAP interpretability. Results: CV-OOF AUCs ranged from 0.697 (LR) to 0.795 (GLMNET), suggesting apparent discriminative ability. However, independent test-set AUCs collapsed to 0.467-0.574, with all 95% confidence intervals including 0.5. Calibration slopes were near zero or negative (range: -0.219 to 0.102), and decision curves revealed negligible clinical net benefit beyond the "treat all" strategy. The striking discrepancy between CV-OOF and test performance indicates that the models failed to generalize. SHAP analysis identified candidate probes (CCNI, MSMB, YBX3, ABCC5, and others), but these require biological validation. Conclusion: Under a strictly leakage-free evaluation protocol, machine learning models trained on transcriptomic data alone did not predict breast cancer bone metastasis in an independent test set. This negative result has important methodological implications: it demonstrates the importance of independent testing, quantifies the gap between cross-validated and test performance, and cautions against over-optimistic interpretation of cross-validated results in small, imbalanced transcriptomic cohorts. We provide a fully reproducible analytical framework for similar studies.