Abstract / Summary
Abstract Background This study aimed to develop and externally validate a self-supervised deep learning model that integrates admission non-contrast and contrast-enhanced CT for admission risk stratification of Atlanta-defined severe acute pancreatitis (SAP, persistent organ failure > 48 h), thereby addressing limitations of current clinical and imaging-based severity assessments.Affiliations: Kindly check whether the authors affiliations are processed correctly and amend if any. We have thoroughly reviewed all authors’ affiliations. Minor revisions have been made to correct the affiliation of author Shifang Zhou. The updated institutional information is shown in the revised manuscript. Methods This multicenter retrospective study included 2,677 patients with acute pancreatitis from three tertiary centers who were classified as having severe (persistent organ failure > 48 h per the modified Marshall score) or non-severe disease according to the 2012 revised Atlanta Classification. All patients underwent both unenhanced and contrast-enhanced CT within 24 h of admission. An SE-ResNet-50 backbone was pretrained using contrastive self-supervised learning on 10,740 unlabeled CT images from the training cohort (Center I) and was then fine-tuned for SAP prediction using labeled dual-phase scans. Single-modality and multimodal fusion classifiers were developed. Discrimination was evaluated with the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, positive and negative predictive values, accuracy, and F1 score, with paired AUC differences assessed by patient-level bootstrap resampling. Decision curve analysis was used to quantify net clinical benefit, and Grad-CAM and SHAP analyses provided model interpretability. Results A total of 2,677 patients (median age, 41 years [IQR, 34–53]; 2,049 men) were evaluated and assigned to the training cohort ( n = 1,790), external validation cohort I ( n = 535), and external validation cohort II ( n = 352). The Fusion model achieved AUCs of 0.938, 0.930, and 0.911 in the three cohorts, respectively. Its discrimination was statistically equivalent to that of the CECT-only model, indicating that CECT alone carried the dominant predictive signal (AUC range, 0.911–0.938) and was higher than that of the NECT-only model (AUC range, 0.757–0.807; all p < 0.001). Across cohorts, the Fusion model maintained high sensitivity (up to 94.9%) and favorable net clinical benefit. Youden index thresholds for NECT, CECT, and Fusion (0.166, 0.189, and 0.247) were selected in the training cohort and applied unchanged to both external cohorts. A self-supervised CT framework provided effective admission risk stratification for Atlanta-defined SAP, with CECT-derived features carrying the dominant predictive signal. The model identifies patients at elevated risk of persistent organ failure from routine admission imaging and may support standardized triage and monitoring decisions. The NECT-only variant retained practical value when CECT was contraindicated. Because baseline organ-function status was not available, the model may partly reflect early systemic severity already present at admission rather than purely predicting future deterioration; prospective validation with structured clinical endpoints is required.