Abstract / Summary
Atrial fibrillation (AFib) is the most prevalent sustained cardiac arrhythmia, yet most deep learning detectors are validated on a single cohort and few report uncertainty. We developed a compact 14M-parameter convolutional neural network for binary AFib classification from 10 s, 12-lead ECG at 500 Hz, trained exclusively on the HDXML dataset (67,432 ECGs; 9628 AFib) with on-the-fly augmentation for noise, drift and missing channels. Using one unchanged checkpoint, we evaluated six independently sourced public cohorts totaling over 1.25 million recordings, including MIMIC-IV critical care and two ambulatory Holter cohorts. All metrics are reported with 95% confidence intervals from a cluster bootstrap at the patient, recording or subject levels. ROC–AUC exceeded 0.95 in every cohort, from 0.960 (95% CI 0.956–0.963) on CODE-15 to 0.999 (0.998–0.999) on SPH, including 0.966 (0.965–0.967) on MIMIC-IV with 75.3% sensitivity and 97.7% specificity, 0.963 (0.929–0.987) on CPSC 2021 and 0.995 (0.988–0.999) on MIT-BIH. Performance was stable under progressive lead masking without retraining, and median inference took 110 ms per segment on a server-class CPU without a GPU. A compact network trained on one curated source can therefore generalize across hospitals, countries and recording hardware. This study is retrospective; prospective validation is required before clinical use.