Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Africa Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
mic
Host:
Verse-aligned text for 792+ African languages, plus several world languages, for building parallel and monolingual corpora. Every language is aligned on a shared verse key, so any two languages can be joined into a parallel corpus: African ↔ English (English is the default pair) African ↔ African (e.g. Twi ↔ Yoruba, Hausa ↔ Amharic) African ↔ other language (French, Arabic, Chinese, Portuguese) Monolingual corpus for any single language

Visit

huggingface.co

Tasks

machine translation

Languages

AfrikaansAmharicChichewaGandaHausaIgboKinyarwandaLingalaMalagasyOromo+12

Tags

africaafrican-languageslow-resourceparallel-corpusmachine-translationbible

Licenses

other

Similar

AfriSpeech/africa-corpus-builderdatawise-africa/sheria-corpus-v1East Africa: swahili historical corpus pdCorpus-based research on English in AfricaChallenges of Doing Corpus Linguistics in AfricaA Corpus of Illuminated Qurʾāns from Coastal East Africa

AfriSpeech/africa-corpus-builder

Get access to monolingual and parallel data for 693 African languages # Africa Corpus Builder A to

datawise-africa/sheria-corpus-v1

The Sheria Corpus v1 is a curated collection of Kenyan legal case summaries from both the High Court

East Africa: swahili historical corpus pd

pretty_name: Swahili Historical Corpus — Public Domain

Corpus-based research on English in Africa

Abstract This chapter provides linguists and students not yet familiar with corpus-based research

Challenges of Doing Corpus Linguistics in Africa

Corpus linguistics, although a relatively new field of research endeavour, has made great strides in

A Corpus of Illuminated Qurʾāns from Coastal East Africa

Abstract This article examines a little-known corpus of illuminated Qurʾān manuscripts that we