Abstract / Summary
Abstract This study presents a deep learning framework for four-class gastrointestinal (GI) endoscopy risk stratification beyond conventional binary lesion detection. Im- ages from the HyperKvasir dataset were categorized into four clinically motivated risk groups: Normal, Inflammatory, Pre-malignant, and High-Risk, informed by rec- ommendations from the American College of Gastroenterology and European Soci- ety of Gastrointestinal Endoscopy. We introduce an Asymmetric Endoscopy Loss (AEL) function that enables tunable class weighting, allowing greater emphasis on High-Risk cases while balancing overall predictive performance with safety-oriented sensitivity. Three lightweight architectures (DenseNet-121, EfficientNet-B0, and DeiT-Tiny) were evaluated using AEL, with Monte Carlo Dropout for uncertainty estimation and a referral strategy that escalates Pre-malignant and High-Risk pre- dictions for further review regardless of model confidence. GradCAM was used to provide visual interpretability. DenseNet-121 achieved a macro F1-score of 0.8419 and an AUC of 0.9730. No ground-truth High-Risk images were predicted as either of the two lower-risk tiers on the internal test set, and 43.1 % of images were eligi- ble for simulated automatic clearance under the predefined referral rule. Five-fold cross-validation provided an additional assessment of performance stability, while McNemar tests quantified paired differences between architectures. These findings support further investigation of tunable risk-weighted losses, lightweight architec- tures, and uncertainty-aware referral strategies for developing interpretable and safety-oriented GI endoscopy decision-support systems.