Abstract / Summary
Background: The Greater Manchester Care Record (GMCR) integrates health and social care data for 2.8 million individuals. While it contains maternity information, identifying individual pregnancies, including timing and outcomes, is complex due to variations in healthcare interactions and recording practices. This study develops open-source code to create a pregnancy register within the GMCR using diagnosis- and procedure-based algorithm established in other databases. Methods: 1,448 maternity SNOMED clinical codes identified pregnancy-related primary care records among women aged 14-49 years between March 2012 and August 2023. Codes were grouped by pregnancy outcome, with a hierarchical ruling approach where multiple outcomes appeared within one episode. Pregnancy start dates were derived from hierarchical coded events or imputed using biologically plausible duration based on outcome type. Internal validation compared GMCR register episodes against Hospital Episode Statistics (HES) hospital maternity admissions; external validation compared outcome rates against published UK estimates. Results: The algorithm was developed across 379,832 women, identifying 880,926 pregnancies. Internal validation showed nearly half of register deliveries had a corresponding HES record, with a positive predictive value (PPV) of 48.1%and a median date difference of 17 days (IQR 1-30). External comparison revealed live birth (927.8 vs. 995.6 per 1,000 deliveries) and stillbirth (2.12 vs. 3.9 per 1,000 births) rates close to national estimates but a notable discrepancy for miscarriage (8.00 vs. 15.3 per 100 pregnancies) and termination of pregnancy (5.36 vs. 25.1 per 1,000 women). Demographic alignment was strong for maternal age (31 vs. 29.8 years) and deprivation (IMD10 median: 3 vs. 2.6), though ethnic minority representation was expectedly higher in the GMCR (37.0% vs. 28.7%). Conclusions: This pregnancy register enriches the GMCR for maternal health research, enabling more comprehensive and regionally tailored studies to inform healthcare delivery and policy. This work represents an important first step in developing a scalable, open-source algorithmic tool to support pregnancy research across diverse populations. By making the code openly available, it can be adapted for use in federated data environments, enhancing reproducibility and collaboration.