
Background: Preterm birth (PTB) is a leading cause of neonatal morbidity and mortality, particularly in low- and middle-income countries. Existing clinical risk assessment tools are limited in their ability to provide early and generalizable prediction using routinely collected antenatal data.
Objective: This study aimed to develop and validate an explainable machine learning (ML) model for PTB risk prediction using routine antenatal care (ANC) data from multiple health facilities in Ethiopia.
Methods: A retrospective multicenter dataset comprising 1,356 maternal records from three health facilities was used. A binary risk stratification (BRS) framework was developed using eight state-of-the-art (SOTA) ML models, including tree-based and tabular deep learning (DL) approaches. The best-performing model was further adapted into a one-month-ahead prediction (OMHP) model by excluding variables occurring close to delivery to preserve temporal validity. Model performance was evaluated using discrimination (area under the receiver operating characteristic curve [AUC]), calibration (Brier score), and decision curve analysis (DCA). External validation was performed using an independent dataset from a separate health facility. Model interpretability was assessed using SHapley Additive exPlanations (SHAP) and Local Interpretable Model-Agnostic Explanations (LIME).
Results: The dataset included 402 PTB cases (29.6%). The TabPFN-based BRS model achieved an AUC of 0.976 (95% CI 0.958 - 0.993) and a Brier score of 0.050. The OMHP model achieved an AUC of 0.959 (95% CI 0.929 - 0.984) and a Brier score of 0.1376. On external validation, the BRS model achieved an AUC of 0.90 (95% CI 0.845 - 0.947), while the OMHP model achieved an AUC of 0.872 (95% CI 0.810 - 0.925). SHAP and LIME analyses identified prior PTB history, hypertensive disorders, and ANC visit frequency as the most influential predictors. Both models demonstrated positive net benefit across a range of clinically relevant threshold probabilities in DCA.
Conclusion: The proposed explainable ML models demonstrated good discrimination and calibration for PTB risk prediction using routine ANC data in a low-resource setting. The results suggest potential utility for early risk stratification; however, prospective validation in diverse populations is required before clinical implementation.