Abstract / Summary
Abstract Dementia and its earlier stages mild cognitive impairment (MCI) and subjective cognitive decline (SCD) are increasingly common, yet automated tools that distinguish all four cognitive states simultaneously remain rare. Most published models address a simpler binary problem; the four-class case, which is what clinicians actually face, has attracted far less attention. We developed and benchmarked twelve machine learning classifiers spanning linear, kernel-based, tree-based, boosting, probabilistic, distance-based, and ensemble meta-learning paradigms for four-class cognitive stratification (Cognitively Normal, SCD, MCI, Dementia) using ACE-III assessment-derived clinical and demographic data ( n = 140, 28 features). Class imbalance was mitigated through pipeline-integrated Synthetic Minority Over-sampling Technique (SMOTE), with model evaluation conducted via stratified five-fold cross-validation. All performance metrics (accuracy, macro F1-score, ROC-AUC, balanced accuracy, Cohen’s Kappa, Matthews Correlation Coefficient) were computed across folds and reported as mean ± standard deviation. SHAP (SHapley Additive exPlanations) analysis was employed to quantify global and local feature contributions across all model types. LightGBM achieved the highest macro F1-score (0.723 ± 0.114), ROC-AUC (0.944 ± 0.030), and Cohen’s Kappa (0.729 ± 0.138) among the twelve classifiers, with a balanced accuracy of 0.716 ± 0.121. Voting Classifier attained a near-identical balanced accuracy (0.716 ± 0.068) and the highest single-model MCC (0.734 ± 0.091), closely matching LightGBM’s own MCC (0.733 ± 0.139), and ranked second by macro F1-score (0.714 ± 0.058). k-NN and SVM ranked third and fourth by macro F1-score (0.640 ± 0.136 and 0.631 ± 0.095, respectively), while Gaussian Naïve Bayes performed least favourably (F1: 0.480 ± 0.143). A paired t-test found LightGBM’s macro F1-score advantage statistically significant against one of the eleven remaining classifiers, Logistic Regression (t = 6.783, p = 0.0025, Holm-adjusted p = 0.0271), while differences from the other ten classifiers — including its closest competitors, Voting Classifier, k-NN, and Extra Trees — did not reach significance at this sample size. SHAP analysis identified age, health-condition status, educational attainment, depression, blood pressure, and family history of MCI/Alzheimer’s disease as the most consistently influential predictors across model types. LightGBM classifier trained on Egyptian Arabic ACE-III data classifies patients into four cognitive categories with accuracy that is promising for pre-screening or triage role, though external validation on independent cohorts is needed before any claim of clinical utility can be made, and SHAP makes its reasoning transparent. The framework requires no specialist equipment and runs on data already collected in routine Egyptian encounters, making primary care deployment in Egypt and the MENA region practical pending that validation.