Abstract / Summary
Alzheimer’s disease MRI classification remains challenging because differences between adjacent cognitive-severity categories can be visually subtle and affected by class imbalance, heterogeneous atrophy patterns, limited frequency-domain representation, and unreliable predictive confidence. This study proposes MS-AtroNet, a multi-stream deep learning framework for four-class image-level MRI slice classification into NonDemented, VeryMild, Mild, and Moderate categories. The architecture combines a four-stage convolutional backbone with Multi-Receptive Field Attention Blocks for local, anisotropic, and dilated contextual feature extraction. A VMamba-S stream processes the deepest Stage 4 representation for long-range contextual modeling, while a one-level Haar discrete wavelet transform extracts LL, LH, HL, and HH subband information for frequency-aware modulation. Cross-Stream Attention Fusion integrates convolutional queries with VMamba-enhanced keys and values under wavelet-derived gating, and a Hierarchical Feature Pyramid aggregates Stage 2, Stage 3, and fused Stage 4 representations. An Evidential Classification Head produces class probabilities and image-level uncertainty estimates. Experiments used 15,270 augmented training images, 1,999 original validation images, and 3917 original internal test images from the Kaggle Augmented Alzheimer’s MRI Dataset. MS-AtroNet achieved an internal-test accuracy of 0.978, Macro-AUC of 0.993, AUC-PR of 0.957, Macro-Se of 0.957, Macro-Sp of 0.990, and Macro-F1 of 0.961 with 38.4 million trainable parameters. Calibration results included an ECE of 0.018, MCE of 0.061, and Brier Score of 0.041. The model also showed the smallest mean Macro-F1 degradation under the evaluated controlled synthetic perturbations, at − 0.014. Validation-derived uncertainty thresholds improved retained-case performance as uncertain predictions were excluded. Exploratory evaluation on 756 Kaggle ImageOASIS images achieved an accuracy of 0.903, Macro-AUC of 0.913, and Macro-F1 of 0.867. These findings support promising image-level performance, while patient-disjoint, multisite, prospective, and clinically validated evaluation remains necessary.