Abstract / Summary
Abstract Background Clinical prediction models are increasingly considered for decision support in emergency care, but high discrimination alone does not establish temporal reliability. We evaluated whether models based only on data available at initial emergency department assessment retained calibrated and threshold-specific performance over time and across clinically important subgroups. Methods We conducted a retrospective multicentre cohort study of 626,233 adults with acute trauma recorded in the Korean Emergency Department-based Injury In-depth Surveillance database from 2019 to 2023. Logistic regression, random forest, Extreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM) models using 15 initial-assessment variables were developed in 2019–2022 and compared in a temporally separated 2023 cohort. Evaluation included AUROC, AUPRC, Brier score, probability calibration, derivation-defined operating thresholds, decision-curve analysis, and age- and injury-mechanism subgroup performance. Because 2023 performance informed designation of XGBoost for extended characterisation, the 2023 cohort was not treated as independent external validation of a prespecified final model. Results In 2023, 1,653 of 119,234 patients (1.39%) died in hospital. AUROCs were 0.977 for logistic regression, 0.980 for random forest, 0.978 for XGBoost, and 0.978 for LightGBM; corresponding AUPRCs were 0.630, 0.708, 0.714, and 0.716. For XGBoost, derivation-fitted isotonic calibration improved the calibration slope from 0.919 to 1.021 and reduced the Brier score from 0.0367 to 0.0065. A sensitivity-constrained operating threshold that achieved 90.01% sensitivity during derivation achieved 79.49% when applied unchanged in 2023. AUROC was 0.993 among adults aged 19–64 years and 0.930 among adults aged 65 years or older. Conclusions Routinely collected initial emergency department data carried substantial prognostic information, but strong aggregate discrimination coexisted with temporal instability of fixed operating points and lower discrimination in older adults. These findings support distinguishing ranking performance, calibrated risk estimation, and threshold-specific classification when prediction models are evaluated for clinical decision-support use, with attention to temporal transportability and subgroup reliability. Independent validation is required before clinical implementation.