Abstract / Summary
Background/Objectives: Acute appendicitis is the most common surgical emergency of childhood, yet its diagnosis remains difficult because presentations are atypical, inflammatory markers are nonspecific, and the consequences of error run in both directions, from negative appendectomy to missed perforation. Over the past two decades, a large body of work has tried to support this decision with clinical scores, and more recently with machine learning, deep learning, and large language models. Methods: This critical methodological review synthesises that literature across the whole care pathway rather than the single question of binary diagnosis, using a reproducible search and a transparent study-level appraisal, and it therefore covers severity stratification, prediction of non-operative treatment response, imaging stewardship, and the postoperative course. We evaluate this evidence through the lens of contemporary methodological standards, in particular the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis, updated for artificial intelligence (TRIPOD+AI), and the Prediction model Risk Of Bias Assessment Tool, updated for artificial intelligence (PROBAST+AI). Results: We describe the main families of models, from the Alvarado and Pediatric Appendicitis scores to random forests, gradient boosting, convolutional networks applied to ultrasound, and emerging generative models. Reported performance is dominated by discrimination while calibration, clinical utility, external validation, fairness, and reproducibility are reported inconsistently and are often absent. We argue that the recurring pattern of very high reported accuracy reflects methodological fragility more than clinical readiness. Conclusions: We offer a practical checklist for the critical appraisal of appendicitis prediction studies together with a research agenda aimed at closing the gap between a high area under the curve and safe use at the bedside.