Abstract / Summary
Abstract Background Pulmonary embolism (PE) is a life‑threatening cardiovascular emergency with non‑specific presentations, making timely diagnosis challenging. Over‑reliance on computed tomography pulmonary angiography (CTPA) has raised concerns about radiation exposure and healthcare costs. Traditional clinical prediction rules, such as the Wells and Geneva scores, have limited specificity and fail to capture complex interactions among risk factors. Machine learning (ML) offers a promising alternative, but many existing models lack interpretability or depend on features not routinely available in emergency settings. Objective To develop and internally validate a parsimonious, interpretable ML model for predicting PE using routinely available clinical variables, and to apply SHapley Additive exPlanations (SHAP) to enhance model transparency. Methods We conducted a retrospective study of 519 patients suspected of PE at a single tertiary referral center. Eight independent predictors—age, hyperlipidemia, limb circumference, D‑dimer, oxygen saturation, dyspnea, immobility, and malignancy—were identified via univariate and multivariable logistic regression on the training set. Three ML algorithms (Lasso, Random Forest, and XGBoost) were trained and evaluated on a 70/30 train‑test split. Model performance was assessed using AUC‑ROC, calibration plots, and Brier score. SHAP analysis was performed for global and local interpretability. Results On the testing set, Lasso achieved an AUC of 0.766 (95% CI: 0.691–0.840), while XGBoost showed a comparable AUC of 0.764 (95% CI: 0.690–0.837). XGBoost was selected as the primary model due to its superior interpretability via SHAP. At the optimal threshold of 0.526, XGBoost achieved a sensitivity of 74.4%, specificity of 66.7%, and overall accuracy of 71.0%. Calibration was moderate (Brier score = 0.224 vs. null model 0.247; Hosmer‑Lemeshow P = 0.563). SHAP analysis identified limb circumference, D‑dimer, and malignancy as the top three contributors to PE risk. Conclusions The XGBoost model based on eight routine clinical features demonstrated moderate discriminative ability and good interpretability for predicting PE. This tool may serve as a decision‑support aid for emergency physicians in risk stratification, potentially reducing unnecessary CTPA utilization. However, the high PE prevalence (55.6%) in our tertiary cohort limits generalizability; external validation in broader, low‑prevalence populations is essential before clinical implementation. Trial registration : Not applicable (retrospective observational study).