Abstract / Summary
Abstract A binary endoscopy-image classifier can label an image Normal even when the supplied reference label is Celiac. We tested whether technical instability and simple quality measurements improve the prioritization of such disagreements for human review beyond model confidence alone. We audited a supplied open dataset of 188 JPEG images (89 Normal, 75 Celiac, 24 Doubtful). After excluding Doubtful images and five exact duplicate copies, 159 distinct images entered a prespecified five-fold, group-disjoint ResNet-18 experiment. Every image received one held-out prediction. Among predicted-Normal images, we compared a fixed equal-rank combination of weak Normal confidence, prediction changes after mild brightness/JPEG perturbations, and blur/exposure/glare measured in a central image window with confidence alone at matched review rates. The model correctly classified 122/159 images (76.7%); 73 were predicted Normal, including 13 Celiac-labeled images. At the primary 20% review rate (15 images), both methods caught 3/13 disagreements (paired difference 0; grouped-bootstrap 95% interval, -3 to 3). At 10%, confidence caught 2 versus 1; at 30%, both caught 5. In a prespecified central-field training-crop analysis, the combined score caught 4/12 versus 6/12 for confidence at 20% (difference -2; bootstrap interval, -5 to 2). The added signals did not demonstrate better false-normal image-label capture at equal review workload. The finding is a leakage-aware negative test, not evidence about missed clinical diagnoses or deployment safety.