Abstract
Background
Child stunting remains a critical public health challenge in Ethiopia. Standard parametric frameworks fail to capture complex non-linear interactions among multidimensional determinants. This study evaluated six supervised machine learning (ML) algorithms, incorporated probability calibration and performed clinical utility diagnostics to predict under-five child stunting.
Methods
We analyzed secondary data from 2024/25 the Ethiopian Demographic and Health Survey (EDHS) Kids Recode dataset (N = 10641). Six algorithms Logistic Regression, Decision Tree, Support Vector Machine, Random Forest, Gradient Boosting and XGBoost were trained on 80% of the dataset (n = 8,513) using 10-fold cross-validation and tested on a 20% holdout sample (n = 2128). Isotonic probability calibration was implemented for tree ensembles. Models were benchmarked using discrimination metrics, calibration fit and Decision Curve Analysis.
Results
National stunting prevalence was 36.8% (95% CI: 35.9%–37.7%). Non-parametric ensemble models outperformed linear models. Post calibration Gradient Boosting achieved an optimized AUC-ROC of 0.892 (95% CI: 0.868–0.916) with superior calibration (Hosmer-Lemeshow p > 0.05). At a p = 0.50 decision threshold, operational metrics reached 85.1% Accuracy, 82.4% Precision, 83.6% Sensitivity, 86.0% Specificity and an 83.0% F1-score. Primary predictive drivers included maternal BMI, child age gradient, household wealth quintile, maternal educational attainment and recent diarrheal illness.
Conclusion
Isotonic-calibrated Gradient Boosting provides highly accurate, reliably calibrated risk stratification for child stunting. Integrating these algorithmic decision support tools into community digital platforms (e.g. Ethiopia's Health Extension Program) enables proactive targeted screening before irreversible growth failure occurs.