AfriCorpus-v1 is the first public release of LocaleNLP's audited, deduplicated, and quality-filtered African language corpus. Built to power the AfriLION LLM project, this dataset directly addresses the Tokenizer Fertility problem that causes all current LLMs to underperform on African languages.
Language
Code
Script
CC-100 Source
Status
Wolof
wo
Latin
CC-100
Audited
Swahili
sw
Latin
CC-100
Audited
Hausa
ha
Latin + Ajami