Abstract / Summary
Open-source deep learning models promise scalable quantification of MRI-visible perivascular spaces (PVS), an emerging marker of glymphatic dysfunction in Alzheimer's disease (AD). Whether these models remain valid when applied to heterogeneous AD cohorts and preserve ground-truth derived biological associations has not been established. We benchmarked five open-source nnU-Net-based PVS segmentation models (nnU-Net T1, nnU-Net T1+FLAIR, mcPVS-Net, MedNet-PVS, ADNI-PVS-Net) against manual labels in 60 multi-site ADNI-3 participants balanced across cognitively unimpaired, mild cognitive impairment, and AD dementia, and in 12 publicly available scans from the VAscular Lesions object-level and segmentatiOn (VALDO) challenge. We evaluated overall, voxel, object, and lesion-wise Dice; stratified performance by region, diagnosis, scanner, PVS severity, and white matter hyperintensity (WMH) burden; and tested whether model-derived PVS burden reproduced ground-truth derived associations with AD-related biomarkers. Despite high internal performance, all models showed a specialist-generalist trade-off: the AD-trained ADNI-PVS-Net achieved highest internal performance (overall Dice = 0.58 {+/-} 0.14) but degraded most externally (Dice [≤] 0.11), while MedNet-PVS and mcPVS-Net generalized externally (Dice 0.27-0.30) but systematically misestimated PVS burden. Models were relatively robust across scanner and diagnosis but inconsistent across PVS severity, and some misclassified up to 32% of WMH as PVS. Notably, only AD-trained models reproduced the expected PVS-amyloid biomarker relationship; other models returned a null result, despite comparable or superior Dice scores. These findings show that internal segmentation accuracy and external generalizability alone do not guarantee downstream inferential validity. Careful model selection and cohort-specific validation against manual labels should precede clinical inference in AD.