Abstract / Summary
Background: Community health check-ups routinely collect non-cognitive data. Whether these support mild cognitive impairment (MCI) triage, and whether machine learning outperforms regression, is unclear. Methods: Cross-sectional analysis of 2650 community-dwelling adults aged ≥65 in Akita, Japan (2018–2025). MCI was defined by the NCGG-FAT (1.5 SD, age/education-adjusted norms). Thirty-four pre-specified non-cognitive predictors were modelled using tree ensembles, penalised regression, other machine-learning families and simple reference scores, with stratified 70/30 splitting, DeLong tests, 100 repeated splits, training-derived thresholds, calibration metrics and decision-curve analysis. Results: MCI prevalence was 28.6%. Random forest discriminated best (ROC AUC 0.722, 95% CI 0.681–0.761; PR-AUC 0.542) and was well calibrated (slope 1.048; intercept −0.024), but its advantage over LASSO regression was marginal (ΔAUC +0.031, p = 0.044) and absent versus gradient boosting (p = 0.059). All multivariable models greatly outperformed the Kihon Checklist (AUC 0.550, p < 0.001). Neither operating point was deployable: at 0.5, sensitivity was 0.269 (166/227 cases missed); at the training-derived Youden threshold (0.321), sensitivity was 0.604, specificity 0.699, and PPV 0.445. Net benefit exceeded assess-all/assess-none between threshold probabilities 0.15–0.60. Conclusions: Routine non-cognitive data support moderate, well-calibrated MCI discrimination, but machine learning added little over regression. This is an exploratory triage aid requiring external validation.