Abstract / Summary
Abstract Background: Most deaths from acute pancreatitis (AP) occur in the one in five patients who develop persistent organ failure, that is severe acute pancreatitis (SAP), and existing severity scores discriminate modestly. We developed and internally validated a model for SAP from routinely available first-24-hour data and benchmarked it against machine learning. Methods: Retrospective cohort study using MIMIC-IV version 2.2. Adults admitted to an intensive care unit (ICU) with a first hospitalisation for AP (ICD-9 577.0 or ICD-10 K85.x) and an ICU stay of at least 24 hours were included. SAP was defined per the revised Atlanta classification as persistent organ failure beyond 48 hours, operationalised with a daily modified Marshall score. 62 first-24-hour candidate predictors were considered. Missing data were handled by multiple imputation by chained equations, excluding the outcome. Predictors were selected by least absolute shrinkage and selection operator (LASSO) within each imputation, retaining variables selected in at least three of five; a multivariable logistic model was pooled by Rubin's rules and presented as a nomogram. Performance was assessed by 10-fold cross-validation and bootstrap, with calibration and decision curve analysis, and benchmarked against eight machine learning algorithms. Reporting followed TRIPOD+AI. Results: Of 5,881 AP hospitalisations, 734 patients formed the analytic cohort and 290 (39.5%) developed SAP; hospital mortality was 23.4% versus 6.1%. LASSO selected 21 predictors. The model achieved a cross-validated area under the receiver operating characteristic curve of 0.877 (95% CI 0.852-0.903), a calibration slope of 0.884 (95% CI 0.755-1.013) and a Brier score of 0.136; bootstrap-corrected AUC was 0.889. Discrimination exceeded the modified BISAP score (AUC 0.693; p<0.001), with a net reclassification improvement of 0.600. Across eight machine learning models (test AUC 0.818-0.850), none of which significantly outperformed logistic regression (p=0.494). SHAP ranked minimum Glasgow Coma Scale, creatinine, mechanical ventilation and blood urea nitrogen highest. Removing organ-support variables reduced AUC to 0.816 (p=0.016); adding five sparse biomarkers did not improve it (p=0.143). Conclusions: A first-24-hour nomogram predicted severe acute pancreatitis with good discrimination and calibration, matching eight machine learning models while remaining transparent and usable at the bedside. External validation is required before clinical adoption.