Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriCorpus v1

Domain:

natural language processing

Record type:

dataset
Creator:
Loc
Host:
AfriCorpus-v1 is the first public release of LocaleNLP's audited, deduplicated, and quality-filtered African language corpus. Built to power the AfriLION LLM project, this dataset directly addresses the Tokenizer Fertility problem that causes all current LLMs to underperform on African languages. Language Code Script CC-100 Source Status Wolof wo Latin CC-100 Audited Swahili sw Latin CC-100 Audited Hausa ha Latin + Ajami

Visit

huggingface.co

Tasks

language modeling

Languages

AmharicHausaIgboSomaliSwahiliTigrignaWolofYorubaZulu

Tags

african-languagesnlpmultilingualtext-generationlow-resource

Licenses

cc-by-4.0

Similar

NjugunaKelvin/africorpus-coreSomaliWeb v1dagm v1tinkvu/amharic-v1neh7777/Pretraining-V1rcwhytock/SmartCams: V1

NjugunaKelvin/africorpus-core

A data infrastructure initiative for african languages # Africorpus Core **A data engineering fram

SomaliWeb v1

📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokeni

dagm v1

Introduction: Dysphagia is the common post-stroke complication. It is a common cause of prolonged ho

tinkvu/amharic-v1

neh7777/Pretraining-V1

A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This

rcwhytock/SmartCams: V1

Real-time alerts from AI-enabled camera traps using the Iridium satellite network: a case-study in G