Abstract / Summary
Identifying progressive dementia during its prodromal stage is a challenging yet critical task for healthcare providers. Traditional clinical assessments are often not sensitive enough to detect warning signs of cognitive decline. Machine learning (ML) methods have shown promise in modeling clinical trajectories associated with neurodegenerative disorders such as Alzheimer's disease (AD). Validation of these models' forecasting abilities, however, is widely lacking in the literature. To address this gap, we established a longitudinal cohort from the National Alzheimer's Coordinating Center (n = 15,727) comprising individuals who were cognitively normal (CN) or had mild cognitive impairment (MCI) at baseline. Cognitive status and clinical observations collected over up to 10 visits were used to train several ML classifiers. We employed the gradient-boosted tree ensemble method XGBoost with a custom feature-engineering pipeline. On non-imputed data, our models achieved an AUC-ROC of 0.9447 [0.9417-0.9476] and PR-AUC of 0.8554 [0.8497-0.8611] for the CN to MCI task, and an AUC-ROC of 0.9059 [0.8983-0.9134] and PR-AUC of 0.9214 [0.9156-0.9272] for the MCI to AD task. In a lead-time forecasting analysis spanning mean follow-up periods of 56.1 months (range, 9.1-214.1) for CN to MCI and 42.7 months (range, 9.1-126.1) for MCI to AD, the non-imputed models detected future progression before diagnosis in 29.58% [28.29%-30.87%] of CN to MCI progressors and 22.41% [21.36%-23.46%] of MCI to AD progressors. Among these individuals, the mean lead time was 33.04 months [32.52-33.56] and 17.95 months [17.71-18.19], respectively. False alarm rates among non-progressors were 12.6% [11.5%-13.7%] and 20.0% [18.4%-21.6%], respectively. These findings suggest that longitudinal machine learning models can identify future cognitive decline before diagnosis, indicating clinical utility for early dementia risk stratification.