Background: Diabetes mellitus is rising rapidly across low- and middle-income countries, yet early detection in Sub-Saharan Africa (SSA) remains limited. Context-specific, non-laboratory risk-screening tools could support community case finding. We aimed to develop and validate simplified diabetes risk-prediction models for SSA, with some focus on Malawi and young adult populations.
Methods: We pooled WHO STEPwise (STEPS) survey data from nine African countries totalling 47,594 participants. After harmonisation though, complete-case analytic sample comprised 31,625 adults aged 18-69 years. Eighteen models ranging from logistic and penalised regression, discriminant analysis, tree-based methods, k-nearest neighbours, support-vector machines and a neural network were trained on a stratified 70/15/15 split with five-fold cross-validation. The primary metric was the area under the ROC curve (AUC).
Findings: Pooled diabetes prevalence was 11.0%. Gradient boosting achieved the best discrimination (test AUC = 0.701), seconded by logistic regression (0.698). Support-vector machines (0.518) and k-nearest neighbours (0.653) performed worst. Age and BMI were the dominant predictors, followed by diet, education and salt intake. Performance fell markedly in adults under 40 (best AUC ≈ 0.61). Adding waist circumference and hypertension raised the AUC by 0.11 on the subset with complete clinical data. For Malawi, logistic regression performed best (cross-validated AUC = 0.645).
Interpretation: Simplified, non-laboratory risk models can support diabetes screening in SSA, performing comparably to some established tools. Accuracy was modest especially in younger adults. Incorporating waist circumference, hypertension and family history, alongside external validation, could improve performance and enable integration into community and primary-care screening.