Abstract / Summary
Fine-grained food-image recognition is difficult because visually similar dishes can differ only in small ingredient or texture cues, while images of the same dish vary in plating, viewpoint, illumination, and background. We propose Dynamic Global Context-enhanced CSWin (DGC-CSWin), a shape-preserving extension of CSWin-L that performs late-stage semantic recalibration. The DGC module constructs four independently normalized spatial context descriptors, selects them with input-conditioned head weights, and coordinates additive and multiplicative channel refinement through a joint Softmax gate. The module adds 88,350 parameters and approximately 0.002 GFLOPs. Under a matched 20-epoch, three-seed protocol, DGC-CSWin achieved 85.85% Top-1 and 98.16% Top-5 accuracy on ChineseFoodNet-208, and 81.43% Top-1 and 95.81% Top-5 accuracy on FoodX-251. The corresponding Top-1 gains over protocol-matched CSWin-L were 2.59 and 1.53 percentage points. Routing ablations, Grad-CAM, head-utilization profiles, and paired feature diagnostics indicate that the module preserves the dominant backbone representation while applying selective residual correction. In a 20-image diagnostic, native self-attention maps showed higher normalized entropy and lower first-grid-token concentration for DGC-CSWin than CSWin-L. Single 40-epoch trajectories peaked before epoch 20 on both datasets, with no later improvement. These results support competitive global-context routing as a low-overhead strategy for fine-grained visual recognition.