# Multilingual Health Question Answering in Low-Resource African Languages
**Machine Learning Techniques I — Final Course Project**
Zindi username: `Winston_Russel_Nji`
**Best public leaderboard score: 0.514958**
(ROUGE-1: 0.5019 | ROUGE-L: 0.4275 | LLM Judge: 0.658)
Demo Video | Full Report | Colab Notebook
---
## Overview
This project addresses the Zindi Multilingual Health QA challenge, which requires systems to answer health questions in the same language they were asked, across nine language-country configurations spanning Akan, Amharic, Luganda, Swahili, and English.
Submissions are scored on three metrics: ROUGE-1 F1, ROUGE-L F1, and LLM-as-a-Judge. The central finding of the project was unexpected: **retrieval beat generation on every meaningful metric**. No generative model tested, including fine-tuned mT5, Flan-T5, and NLLB-200, outperformed simply returning the nearest training answer. The final system uses no neural model at inference time at all.
---
## Dataset
| Subset | Language | Country | Training Examples | % of Total |
|---|---|---|---|---|
| Eng_Uga | English | Uganda | 7,624 | 25.6% |
| Aka_Gha | Akan | Ghana | 4,455 | 14.9% |
| Eng_Gha | English | Ghana | 4,443 | 14.9% |
| Eng_Eth | English | Ethiopia | 3,915 | 13.1% |
| Lug_Uga | Luganda | Uganda | 3,383 | 11.3% |
| Eng_Ken | English | Kenya | 2,080 | 7.0% |
| Swa_Ken | Swahili | Kenya | 2,070 | 6.9% |
| Amh_Eth | Amharic | Ethiopia | 1,845 | 6.2% |
**Total:** 29,814 training records, 6,686 validation records, 2,618 test records.
Four of the five languages use Latin script. Amharic uses Ethiopic (Ge'ez), which has no word-boundary whitespace and falls entirely outside the ASCII range. This single difference breaks word-level tokenization for Amharic and motivates the character n-gram approach used throughout the retrieval experiments.
---
## Approach
The project explored two directions.
**Generative approach:** Fine-tune or zero-shot a multilingual seq2seq model (NLLB-200, mT5- …