This publication presents a secondary epidemiological and machine-learning analysis of publicly available Lassa fever surveillance data from Nigeria covering 2018 to 2021. Using the Zenodo dataset deposited by Chioma Dan-Nwafor, the study describes case distribution, symptom patterns, outcomes, case fatality, reporting delays, and geographic concentration of confirmed Lassa fever cases.
The work also evaluates whether routinely collected demographic, clinical, exposure, and testing-context variables can support leakage-aware classification of confirmed Lassa fever cases. To avoid overestimating model performance, laboratory-result fields, final case-classification variables, and outcome-derived predictors were excluded from model development. Model performance was assessed using random validation, forward-time validation, and geographic-holdout validation.
The findings show that confirmed cases were concentrated in known hotspot areas, particularly Edo and Ondo States, with substantial mortality among confirmed cases. Machine-learning performance was moderate and weakened under stricter temporal validation, highlighting the importance of cautious interpretation, transparent validation, and avoiding data leakage in surveillance-based prediction studies.
Overall, this publication contributes a reproducible, clinically cautious analysis of Nigerian Lassa fever surveillance data and demonstrates how machine-learning methods can be applied responsibly to infectious disease datasets without overstating their diagnostic or operational readiness.