Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

LughaGen Multilingual African Language Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Lug
Host:
LughaGen is a curated multilingual corpus for four Kenyan and East African languages: Swahili (sw), Kikuyu (ki), Kamba (kam), and Luo/Dholuo (luo). It was developed as part of the LughaGen research initiative under JHUB Africa, funded by NVIDIA, with the goal of building large language models for low-resource African languages.

Visit

huggingface.co

Tasks

language modeling

Languages

DholuoGikuyuKambaSwahili

Tags

african-languageslow-resourcekenyapretrainingcorpuslughaGenjhub-africanvidia-funded

Licenses

cc-by-sa-4.0