Abstract / Summary
Artificial intelligence (AI) is being applied to obstetric and gynecologic ultrasound at a rapid pace. Current applications fall broadly into four categories. Assistive tools support image acquisition and standardization; diagnostic tools are used to detect existing conditions such as anomalies; predictive tools attempt to estimate the probability of adverse pregnancy outcomes such as preterm birth, preeclampsia and stillbirth; and finally, prognostic tools attempt to forecast longer-term clinical trajectories. The first two categories are verifiable at the point of care by a trained sonographer. The last two are not, and this is where most methodological shortcomings arise. In this review, we describe frequent challenges and shortcomings seen in published predictive AI work in maternal and child health imaging. These include amplification of existing inequalities and bias from retrospective data, data leakage, inflated performance from poor external validation, and the confusion of statistical prediction with clinical diagnosis. We then discuss practical solutions including appropriate study design, structured data collection purpose-built for model development, context-specific metrics such as sensitivity and screen-positive rate, and external validation in multiple independent cohorts. We highlight automated gestational age estimation from ultrasound images as an example of how AI can genuinely advance clinical practice when built on prospective data and evaluated rigorously. Although we use maternal and child health imaging as the illustrative context throughout this review, the principles we describe are not specific to this domain and apply to the development and evaluation of any AI application in health care.