Abstract / Summary
Rationale: Obstructive sleep apnea (OSA) is highly prevalent yet largely under-diagnosed. Current screening strategies rely on effortful questionnaires with modest accuracy, and prior machine learning models often require resource-intensive inputs and lack external validation. Objectives: To develop and validate machine learning models to predict OSA (apnea-hypopnea index [AHI4%][≥]5) and moderate-severe OSA (AHI4%[≥]15) using only routinely collected electronic health record (EHR) data, test cross-system transferability and compare performance with clinical screeners and a Bayesian risk-updating framework. Methods: Three tree-ensemble families (XGBoost, random forests, and histogram-based gradient boosting) were trained on an adult cohort with sleep-study results at Kaiser Permanente (KPSC), using full (106-feature) and minimal (age, sex, body-mass index, race/ethnicity) predictor sets and tuned by nested cross-validation. XGBoost models were selected, internally and temporally validated, and externally validated at University of Kansas Medical Center (KUMC). Measurements and Main Results: On our internal hold-out set (n=53,048), the full XGBoost model reached a ROC-AUC of 0.791 for any OSA and 0.768 for moderate-severe OSA. The minimal model retained most performance (0.764 and 0.736). Temporal performance was stable. Externally, frozen models showed lower AUCs (0.718 and 0.690 for full models; 0.725 and 0.706 for minimal models) but similar to KUMC-native models and improved further with transfer learning. The models significantly outperformed STOP-BANG (AUC difference 0.06-0.09 for predicting any OSA; p<0.001). Bayesian updating with select patient-reported items further improved discrimination. Conclusions: Machine learning models based on EHR data can predict OSA probability that outperform existing screeners, transfers across systems, and offers potential for EHR implementation.