Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Variable-Wise Hybrid Imputation Framework Using kNN, MissForest and PMM for Enhancing HIV Survey Data Quality in Kenya

Domain:

healthcare

Record type:

paper
Creator:
MacKerToo
Editor:
Dep
Publisher:
CCSD
Host:avatar
International audience Missing data remains a critical challenge in large-scale public health datasets, particularly in HIV surveillance, where incomplete observations can bias estimates and weaken decision making. This study proposes a variablewise hybrid imputation framework that integrates k- Nearest Neighbors (kNN), MissForest, and a Modified Predictive Mean Matching (Modified PMM) under the Missing at Random (MAR) assumption. The method employs a composite scoring function to dynamically select optimal donors for each missing observation by combining structural similarity, predictive alignment, and model-based deviation. The framework was applied to HIV survey data (imbalance ratio 8.65:1) and evaluated against individual imputation methods usingboth regression and classification as well as the distributional imputation quality metrics. The Hybrid approach achieved superior imputation accuracy, with the lowest RMSE (0.4297) and MAE (0.3623). It also demonstrated improved classification performance, achieving the highest accuracy (73.43%), specificity (0.7376), balanced accuracy (0.7201), and F1-score (0.3291). McNemar’s test confirmed statistically significant improvements over Modified PMM (p = 0.041), kNN (p = 0.020), and MissForest (p = 0.045). The Hybrid method further exhibited improved probability calibration, with a lower Expected Calibration Error (ECE = 0.2940). Precision-Recall analysis confirmed the Hybrid framework as the best-performing method under class imbalance, achieving the highest Area Under the Precision-Recall Curve (AUPRC = 0.3709), corresponding to a 4.01× lift over the random classifier baseline. An ablation study confirmed that the full three-component hybrid outperforms all two-component subsets and the equal-weights configuration, establishing that performance gains arise from the composite design rather than any single constituent. These findings highlight the effectiveness of adaptive, observation-level donor selection in improving imputation and downstream predictive performance under class imbalance.

Visit

hal.science

Tags

[MATH]Mathematics [math]

Similar

Variable-wise missing data in test data (Sudan).Variable-wise missing data in test data (PRIEST).Heuristics for the Variable Sized Bin Packing Problem Using a Hybrid P-System and CUDA ArchitectureAn Integrated Multi-Method Framework for Gender-Based Violence Research: A Synthetic Data Demonstration Using Kenya Demographic and Health Survey ParametersAkajiaku11/Hybrid-Machine-Learning-Framework-for-Water-Quality-Assessment-and-Contamination-ClusteringData Sheet 2_Methodological guidance for predictor variable selection for adolescent smoking outcomes in Global Youth Tobacco Survey using R and Python.zip

Variable-wise missing data in test data (Sudan).

COVID-19 infection rates remain high in South Africa. Clinical prediction models may be help

Variable-wise missing data in test data (PRIEST).

COVID-19 infection rates remain high in South Africa. Clinical prediction models may be help

Heuristics for the Variable Sized Bin Packing Problem Using a Hybrid P-System and CUDA Architecture

The Variable Sized Bin Packing Problem has a wide range of application areas including packing, sche

An Integrated Multi-Method Framework for Gender-Based Violence Research: A Synthetic Data Demonstration Using Kenya Demographic and Health Survey Parameters

Abstract Background Current research on Ge

Akajiaku11/Hybrid-Machine-Learning-Framework-for-Water-Quality-Assessment-and-Contamination-Clustering

Hybrid Machine Learning Framework for Water Quality Assessment and Contamination Clustering in the N

Data Sheet 2_Methodological guidance for predictor variable selection for adolescent smoking outcomes in Global Youth Tobacco Survey using R and Python.zip

Background

The Global Youth Tobacco Survey is one of the most important sources of data on adolesc