Abstract / Summary
Abstract Background Machine learning is increasingly used to predict depression, but its advantages over conventional statistical models remain uncertain. Apparent improvements may reflect weak comparators, inadequate validation, or data leakage. This study compared machine learning with additive and flexible logistic regression under a leakage-controlled multicohort simulation framework. Methods Three synthetic datasets represented linear, nonlinear, and heterogeneous predictor–outcome relationships. Each contained 6,000 individuals distributed across six cohorts, yielding 18,000 simulated records. Four cohorts per scenario were used for development and two for held-out evaluation. Six approaches were compared: additive logistic regression, ridge logistic regression, flexible ridge logistic regression, random forest, histogram gradient boosting, and a radial basis function support vector classifier. Preprocessing, model selection, and calibration were restricted to development data, with nested cohort-aware validation. Performance was assessed using discrimination, calibration, Brier scores, classification measures, and decision curve analysis. Additional analyses examined outcome-derived and post-treatment leakage. Results Machine learning did not consistently improve discrimination beyond regression. Mean held-out cohort AUROC ranged from .765 to .773 in the linear scenario, with additive and ridge logistic regression achieving approximately .773. In the nonlinear scenario, flexible logistic regression achieved the highest mean AUROC (.836), compared with approximately .826 to .827 for the evaluated machine learning approaches. Performance was closely grouped in the heterogeneous scenario, with mean AUROC values of approximately .830 to .832. Calibration assessment identified underprediction under cohort heterogeneity. Outcome-derived leakage produced perfect discrimination in the relevant contaminated analyses, while post-treatment information increased apparent AUROC to approximately .931 to .942. Conclusions Within this synthetic benchmark, complex machine learning algorithms provided no consistent discrimination advantage over appropriately specified regression. Leakage produced substantially larger apparent gains than differences between valid models. These findings support rigorous validation, credible regression comparators, and assessment of calibration and decision utility. Repeated simulations and independent clinical validation are required before drawing conclusions about clinical effectiveness or model equivalence.