Abstract / Summary
Abstract Deep learning models for chest X-ray classification may exploit non-anatomical visual cues, including laterality markers, hospital text stamps, and medical hardware, rather than relying exclusively on pathological features. This study investigates peripheral attribution patterns and the effects of anatomical masking in two deep learning architectures: DenseNet-121 and Vision Transformer (ViT-B/16). Using a balanced, patient-stratified cohort of 10,596 chest X-rays from the NIH ChestX-ray14 dataset, we evaluate classification performance and introduce the Out-of-Anatomy Attribution Ratio (OAAR), a metric quantifying the proportion of model attribution outside a predefined central chest region. Across 200 test images, ViT-B/16 exhibited a higher mean OAAR than DenseNet-121 (0.3865 versus 0.2485). Following training with rectangular anatomical masking, DenseNet-121's AUROC decreased from 0.8833 to 0.8686, while ViT-B/16's AUROC decreased from 0.8625 to 0.7760. These findings suggest that the evaluated transformer architecture is more sensitive to peripheral image information under the studied conditions. However, differences in attribution methodology and the use of fixed rectangular masks limit causal conclusions about shortcut reliance. Further investigation using segmentation-based masking, standardized attribution methods, and external validation is warranted. Keywords: Pneumothorax Detection; Medical Image Analysis; Deep Learning; Vision Transformers; Convolutional Neural Networks; DenseNet-121; Shortcut Learning; Explainable Artificial Intelligence; Chest X-ray Classification. Publication status: Research preprint; not peer reviewed.