Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfroMAFT Corpus: Language Adaptation Corpus for African languages

Domain:

natural language processing

Record type:

datasetmodel
Creator:
David Ifeoluwa AdelaniJesujoba O. Alabi
Publisher:
Zenodo
Host:avatar

Language Adaptation Corpus for 17 African languages, English, French, and Arabic.

We used this corpus to train the following pre-trained language models:

  • AfroXLMR
  • AfriMT5
  • AfriByT5
  • AfriMBART

If you use this corpus, please cite the MAFAND paper and mC4 paper. 

Visit

doi.org

Tasks

language modeling

Licenses

info:eu-repo/semantics/openAccessNon-Commercial Government Licencehttps://github.com/spdx/license-list-XML/blob/master/src/Apache-2.0.xml

Similar

emmanuel2406/web-corpus-for-african-languagesSynthetic Text Corpus for African Language ASRAFRIDOC-MT: Document-level MT Corpus for African LanguagesA Corpus for Berber LanguagesSouth African Language ID CorpusLughaGen Multilingual African Language Corpus

emmanuel2406/web-corpus-for-african-languages

# Web Corpus for african languages Created by Emmanuel Rassou Info Doc can be accessed here **War

Synthetic Text Corpus for African Language ASR

This dataset contains 13,488 synthetic sentences across 10 African languages (Bambara, Chichewa, Hausa, Kanuri, Luo, Nande, Somali, Twi, Wolof, Yoruba) generated using large language models (GPT-4o, GPT-4.5, Claude 3.5 Sonnet, Claude 3.7 Sonnet). Each sentence has

AFRIDOC-MT: Document-level MT Corpus for African Languages

This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-t

A Corpus for Berber Languages

International audience

South African Language ID Corpus

LughaGen Multilingual African Language Corpus

LughaGen is a curated multilingual corpus for four Kenyan and East African languages: Swahili (sw),