Abstract / Summary
Importance: Preeclampsia is a leading cause of maternal morbidity and mortality globally. As real-world data sources are increasingly used for maternal health surveillance and pharmacoepidemiologic studies, valid algorithms are essential for accurately identifying preeclampsia in routinely collected healthcare data.
Objective: To identify and characterize algorithms used to ascertain preeclampsia in population-based healthcare databases, summarize validation metrics, and describe the range of reported prevalence estimates.
Evidence review: Following PRISMA guidelines of 2021, PubMed was searched for studies published from January 1, 2010, to July 1, 2025. Observational studies using reproducible code-based algorithms to identify preeclampsia in routinely collected healthcare data were included. Screening and full-text review were performed in duplicate, with discrepancies resolved by a third reviewer. Data extraction focused on algorithm definitions, coding systems, validation methods, and epidemiologic estimates. A total of 289 studies met inclusion criteria, including 13 validation studies.
Findings: Across included studies, 43.6% used Nordic nationwide registries, 30.1% used claims databases, and 12.1% used electronic health records. ICD-10 O14.x was used in nearly all studies; inclusion of O11.x and O15.x varied. Validation studies consistently demonstrated high specificity (> 90%) but low sensitivity (≤ 80%), reflecting under-ascertainment when relying solely on diagnosis codes. More than 70% of studies did not report start or end points for the time window of ascertainment. Reported prevalence ranged from 0.6% to 6.6% across case definitions and settings. The prevalence of severe preeclampsia ranged from 0.6% to 2.2%. Algorithms including O11.x, O14.x, and O15.x showed improved sensitivity in some settings.
Conclusions and relevance: Algorithms for identifying preeclampsia in real-world data are highly specific but variably sensitive, with substantial heterogeneity in coding definitions and ascertainment windows. Standardized reporting and consistent inclusion of O11.x, O14.x, and O15.x may improve comparability, reproducibility, and the accuracy of epidemiologic estimates across data sources. High-specificity algorithms may be particularly useful for comparative safety studies, whereas potential under-ascertainment should be considered when estimating incidence or prevalence.