Logo Lanfrica

Perbandingan Algoritma Machine Learning dalam Prediksi Hipertensi Berbasis Data Indonesian Family Life Survey (IFLS5)

Domain:

healthcare

Record type:

paper
Creator:
IndAhmAniDew
Publisher:
Fak
Host:
Hypertension remains the leading non-communicable disease burden in Indonesia, with a national prevalence of 31.6% in 2023. Early risk identification through machine learning (ML) presents a cost-effective alternative to clinical screening, particularly in resource-limited settings. This study aimed to develop and compare ML-based hypertension prediction models using nationally representative survey data from the fifth wave of the Indonesian Family Life Survey (IFLS5, 2014). A total of 31,077 adult respondents with complete data were analyzed. Nine predictor variables were used: age, sex, marital status, body mass index (BMI), days of illness, outpatient visits, inpatient days, mobility difficulty, and employment status. Hypertension was defined as systolic blood pressure ≥140 mmHg or diastolic blood pressure ≥90 mmHg, measured from the average of the second and third readings. Four ML algorithms were evaluated: Logistic Regression (LR), Decision Tree (DT), Random Forest (RF), and XGBoost. Class imbalance was addressed using balanced weighting. Model performance was assessed using Accuracy, AUC-ROC, F1-Score, Precision, and Recall. Logistic Regression achieved the best discriminative performance (AUC-ROC = 0.795, Accuracy = 72.6%), followed by Decision Tree (AUC = 0.792), XGBoost (AUC = 0.781), and Random Forest (AUC = 0.761). Feature importance analysis using Gini-based scores revealed that BMI (0.382) and age (0.348) were the strongest predictors. These findings suggest that ML models using population-level survey data can effectively support community-based hypertension screening in Indonesia.