Abstract / Summary
Abstract Brain tumor classification from MRI is essential for treatment planning, but manual interpretation by radiologists is time-consuming and prone to inter-observer variability. Convolutional Neural Networks (CNNs) capture local texture effectively but have a limited receptive field, whereas Vision Transformers (ViTs) model global context but require large training datasets and lack spatial inductive bias. Existing hybrid CNN-ViT studies on this task, however, rarely benchmark against classical baselines, report computational costs, provide dual visual explainability, or address the risk of patient-level data leakage. The paper proposes a hybrid architecture that sequentially combines a pretrained ResNet50 backbone with a lightweight Transformer Encoder head, converting intermediate CNN feature maps into visual tokens to jointly model fine-grained local lesion detail and global anatomical context. The model is evaluated on the 7200 image datasets (glioma, meningioma, pituitary tumor, no tumor) using stratified 5-fold cross-validation with statistical significance testing, and it is compared to a PCA-SVM baseline built on the same CNN features to isolate the Transformer’s contribution. Computational efficiency (parameters, FLOPS, latency, peak memory) and dual-level explainability (Grad-CAM for the CNN branch, Attention Rollout for the Transformer branch) are also assessed. The proposed model achieves 97.40% ± 0.24% mean cross-validation accuracy, significantly outperforming the PCA-SVM baseline (paired t-test, p = 2.42 x 10^-5) while requiring only 37.20M parameters and 9.21 GFLOPs, an overhead of only 13.68M additional parameters and 0.95 additional GFLOPS (an 11.5% increase) over the CNN-only backbone. These findings position the proposed framework as an accurate, computationally efficient, and interpretable solution for automated brain tumor MRI classification.