Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Machine learning⍰based prediction of cardiovascular disease risk in Africa using WHO Stepwise Surveys: 2014-2019

Domain:

healthcare

Record type:

paper
Creator:
WinFatJanOli
Publisher:
ope
Host:
ABSTRACT Introduction Cardiovascular diseases (CVDs) are the leading cause of death globally, with rising burdens in Africa due to ageing populations, lifestyle changes, and poor risk factor control. Conventional risk scores developed in high-income settings often perform poorly in African populations. Machine-learning (ML) approaches offer potential to improve prediction by capturing complex, non-linear interactions among demographic, behavioural, and biological factors. This study applies ML models to WHO STEPS survey data to generate context-specific CVD risk predictions across 12 African countries. Methods We analysed data from 60,294 adults collected in WHO STEPS surveys between 2014 and 2019 across 12 African countries. Three ML models; Elastic Net logistic regression (LASSO), Random Forest (RF), and XGBoost (XGB); were trained to predict self-reported CVD outcomes. Data were split into training (80%) and testing (20%) sets with five-fold cross-validation. Feature selection used the Boruta algorithm, and model performance was assessed via accuracy, sensitivity, specificity, AUC, F1 score, and Brier score. Results Overall CVD prevalence was 5%. Hypertension emerged as the strongest predictor across all models, followed by alcohol-related harm. Tree-based models outperformed regression approaches and conventional clinical scores, with XGBoost achieving the highest discrimination (AUC=0.769), balanced accuracy (0.699), and calibration (Brier score=0.195). Predicted risk trajectories were smoother and more clinically plausible than Framingham or WHO/ISH scores, particularly across age, sex, and hypertension status. LASSO and Random Forest performed moderately, while conventional risk scores showed poor discrimination and marked miscalibration. Conclusion Machine-learning approaches provide accurate, context-specific cardiovascular risk prediction in African populations. By highlighting modifiable risk factors such as hypertension and alcohol-related harm, these models support targeted interventions aligned with WHO PEN, HEARTS, and SBIRT strategies. The African CVD Risk Prediction Tool translates complex data into actionable insights, offering a scalable platform for prevention-focused, equitable cardiovascular care across diverse African settings.

Visit

doi.org

Licenses

http://creativecommons.org/licenses/by-nc-nd/4.0/

Similar

Dataset and codes for Machine Learning-Based Classification of Self-Reported Cardiovascular Disease History in Africa Using Harmonised Multi-Country WHO STEPS Surveys: 2014–2019Machine learning and chronic kidney disease risk predictionCardiovascular disease risk prediction among people living with HIV in Uganda: A comparison of traditional statistical and machine learning modelsInterpretable ensemble machine learning framework for cardiovascular disease prediction using EMR data and large language models in EthiopiaPredicting the prevalence and determinants of harmful alcohol consumption in Africa: Evidence from WHO STEPS surveys, 2014–2019Implementation of cardiovascular disease risk prediction scores in Sub-Saharan Africa: a scoping review

Dataset and codes for Machine Learning-Based Classification of Self-Reported Cardiovascular Disease History in Africa Using Harmonised Multi-Country WHO STEPS Surveys: 2014–2019

This repository contains the data documentation, derived analytical datasets and outputs, and reprod

Machine learning and chronic kidney disease risk prediction

With a prevalence of approximately 10–15% in Africa and a close relationship with other non-communic

Cardiovascular disease risk prediction among people living with HIV in Uganda: A comparison of traditional statistical and machine learning models

Objective: To develop and compare the performance of traditional statistical models and machine lear

Interpretable ensemble machine learning framework for cardiovascular disease prediction using EMR data and large language models in Ethiopia

Cardiovascular diseases (CVDs) are leading causes of morbidity and mortality globally, with a growin

Predicting the prevalence and determinants of harmful alcohol consumption in Africa: Evidence from WHO STEPS surveys, 2014–2019

ABSTRACT Background: Harmful alcohol

Implementation of cardiovascular disease risk prediction scores in Sub-Saharan Africa: a scoping review

Abstract Background Cardiovascular d