Logo Lanfrica

Evaluating Hallucination Patterns in Small Language Models on African Educational Questions

Domaine:

natural language processingeducation

Type de record:

datasetpapersoftware
Créateur:
Ola
Éditeur:
Zenodo
Hôte:avatar

Large language models are increasingly deployed in educational settings across sub-Saharan Africa, yet their reliability on African educational content remains empirically understudied. This paper investigates whether small open-source language models exhibit systematic reliability gaps when answering African educational questions compared to globally represented content.

 

Using DistilGPT-2 (82M parameters), we evaluate 100 hand-curated questions split evenly between globally represented educational topics and African history, governance, geography, and philosophy. Manual annotation using a 3-point hallucination scale reveals that while the model hallucinated on 92% of global questions, every fully hallucinated African response clustered around political figures and regional institutions underrepresented in standard English-language training corpora — including ECOWAS, Nigeria's first president, and Kwame Nkrumah — while African questions with global English-language presence were answered more reliably.

 

We identify three alignment-relevant problems illustrated by these findings: confidently wrong outputs that give users no signal of failure, invisible distributional harm that aggregate benchmarks cannot detect, and evaluation coverage as a safety property, the principle that which failures are visible depends entirely on which populations are represented in the evaluation dataset.

 

The complete dataset, annotation schema, attention visualizations, and reproducible pipeline are publicly available on GitHub. This work builds on AfriLearn Lens (Olayiwola, 2025, DOI: 10.5281/zenodo.19644518), an earlier study of structural bias in AI-driven educational prediction in African contexts.