Abstract
Background
Accurate gestational age (GA) estimation is fundamental for perinatal research and pharmacovigilance but remains challenging in low-resource settings where routine pregnancy records are inconsistent, early ultrasound is limited, and multiple GA estimates often conflict.
This challenge is most acute in large-scale observational studies and pharmacovigilance registries, where the data quality controls available in clinical trials are rarely feasible.
Methods
This study implemented a two-stage approach to derive more reliable pregnancy conception dates using routine pregnancy data from 42 healthcare facilities in Western Kenya. We first applied Isolation Forest to identify and exclude implausible GA estimates, then used a linear mixed-effects model to synthesise the remaining measurements into a single calibrated “best” estimate, accounting for systematic biases and pregnancy-level variation. The model pools information across all pregnancies to generate a calibrated conception date for each individual woman, even when her own measurements are limited.
Results
Integrating multi-modal inputs in each pregnancy synthesised a unified GA estimate, providing a robust dating consensus even in the absence of early ultrasound and variation across dating methods. Anomaly detection revealed data quality was unevenly distributed, with anomaly rates higher among pregnancy complications: 0.5% in term live births versus 5.1% in miscarriages. These discrepancies were primarily driven by data entry errors in antenatal clinic records (46.8%) and last menstrual period (LMP) date recall. The mixed effects model identified varying precision among dating methods. Using early ultrasound as the gold standard (mean gestation: 278.6 days), later ultrasound scans showed progressive bias, increasing from 3.15 days in the second trimester to 6.41 days in the third trimester, but retained high precision (SD: 7.5–8.9 days). Despite this drift, ultrasound markers remained more precise than other methods, such as LMP, Ballard score, foot length and fundal height, which had wider uncertainty spreads (>35 days).
Conclusion
Routine care datasets have the potential to yield reliable GA estimates in the absence of widespread early ultrasound through rigorous data cleaning and statistical modelling. However, translating this approach into real-time clinical practice will require embedding these algorithms into interoperable electronic medical records.