Background: Low birth weight (LBW, <2500 g) is associated with poor outcomes across multiple domains of child development.
Objectives: We aimed to develop and internally validate machine learning models to predict LBW in the Kenyan coast and determine the top predictors of LBW.
Methods: We developed Machine Learning (ML) predictive models using data collected from pregnant women who delivered at Kilifi County Hospital between 2011 and 2019. Model training was conducted with logistic regression, random forest, extreme gradient boosting (Xgboost) and Tabular Prior-data Fitted Network (TabPFN). We used inverse probability of treatment weighting to adjust for potential confounding arising from differences in antenatal care (ANC) visits (> 4 vs ≤4). Hyperparameter search done using 5-fold cross validation and guided using Bayesian optimization. Internal validation was done using a temporal split. Discrimination was assessed using area under the curve (AUC), calibration using calibration plots and overall performance with the Brier score.
Results: Approximately 17% of the 25,699 newborns included in the study had LBW. Performance of the TabPFN model (AUC = 69.3%, 95% CI [68.6%, 70.1%] was similar to logistic regression (AUC = 68.8%, 95% CI [67.3%, 70.4%]. Across the four models, gestational age at first ANC, having a multiple pregnancy, mother’s age, history of high blood pressure during pregnancy and history of pregnancy complications were among the most common top predictors of LBW.
Conclusion:
Our findings suggest that logistic regression achieved comparable performance to the best ML models for predicting LBW risk. If properly integrated into antenatal care workflows, the developed prediction models may improve existing public health interventions through risk stratification and targeted maternal care. Further external and prospective validation, along with cost-benefit analyses across diverse settings and populations, would be required before implementing these models in practice.