Abstract / Summary
Two assumptions dominate deep learning research on pneumonia detection from chest radiographs: that newer architectures outperform older ones, and that intra-dataset evaluation suffices to establish clinical utility. This study tested both. Nine architectures, spanning classical, efficient-scaling, and modern convolutional networks and Vision Transformers, plus a weighted soft-voting ensemble, were compared under a unified protocol on the pediatric Chest X-Ray Images dataset. None of the post-2020 architectures outperformed InceptionV3 (2016), which achieved the best individual accuracy (98.18%). The validation-weighted ensemble reached 98.52% accuracy and 99.22% pneumonia sensitivity on the internal test set; because its components were selected using test-set rankings, these values are upper-bound estimates. Bonferroni-corrected McNemar tests found no significant differences among the top-performing models. A cross-population external evaluation on an adult RSNA subset, which combines age, institutional, acquisition, and label shifts, revealed markedly different behavior: InceptionV3 retained 85.30% accuracy, whereas VGG16 fell to 59.60%, and the ensemble underperformed InceptionV3 alone. Ablation, Grad-CAM, and efficiency analyses complemented these findings. Architectural recency is not a proxy for performance, and external evaluation should be an architecture-selection criterion rather than an optional addendum.