Abstract / Summary
On the public Regensburg Pediatric Appendicitis Dataset, several recent studies report near-ceiling diagnostic performance using radiologist-interpreted ultrasound findings as tabular model inputs, without processing the underlying images. This raises a question worth answering carefully: how much of that performance reflects genuine diagnostic information, and how much reflects the way this dataset's outcome labels were constructed? We defined four modeling arms, an image-only arm (a backbone-matched BiomedCLIP pipeline as the primary comparator, with an originally-planned ResNet18 model retained as a secondary baseline), a clinical/laboratory-only model (XGBoost), a radiologist-findings model (Random Forest, treated as a reference built from expert-recorded features rather than a live expert comparator), and a multimodal model fusing frozen BiomedCLIP image embeddings with clinical/laboratory data, and compared them using paired DeLong tests, bootstrap confidence intervals, TOST equivalence testing, and pooled repeated cross-validation. The multimodal model (pooled AUC 0.894) showed significantly higher discrimination than the backbone-matched image-only (0.789, p less than 0.0001) and clinical-only (0.855, p less than 0.0001) models. Against the radiologist-findings model (0.934), it initially fell short (p = 0.034). We then found that the dataset's own rule for labeling non-surgically-managed patients uses appendix diameter, one of the radiologist-findings predictors, and, separately, the Alvarado and Pediatric Appendicitis Scores, both retained as clinical/multimodal predictors. Removing appendix diameter dropped the radiologist-findings AUC to 0.790; removing the two scores from the multimodal model changed its AUC only marginally, from 0.894 to 0.890, a difference not statistically distinguishable from no change (p = 0.25). Comparing these two corrected models directly, the multimodal model remained significantly higher (0.890 vs. 0.790, p = 6.2 times 10 to the power of minus 6). A supplementary check found that 137 of 578 patients (24 percent) have ultrasound images containing burned-in caliper measurements; on a matched comparison holding the patient set fixed, excluding these images significantly reduced both the image-only and multimodal models' AUC, indicating a second, image-level information pathway that our design cannot fully attribute to on-screen numeric values versus other differences between marked and unmarked views. An exploratory check that combined both corrections, restricting the doubly-corrected multimodal model to annotation-free images on the smaller, matched patient subset, found a smaller, non-significant difference against the diameter-removed radiologist-findings model (p = 0.176), so our main doubly-corrected comparison should not be read as controlling for this second pathway. The multimodal model showed higher discrimination than the unimodal automated baselines on this benchmark, significantly so in pooled repeated cross-validation; its apparent shortfall against radiologist-derived features is highly sensitive to the dataset's outcome-definition procedure, although the exact fraction of that shortfall attributable to label-construction dependence, as opposed to genuine diagnostic information, cannot be determined from these data alone.