Language Adaptation Corpus for 17 African languages, English, French, and Arabic.
We used this corpus to train the following pre-trained language models:
If you use this corpus, please cite the MAFAND paper and mC4 paper.
# Web Corpus for african languages Created by Emmanuel Rassou Info Doc can be accessed here **War
This dataset contains 13,488 synthetic sentences across 10 African languages (Bambara, Chichewa, Hausa, Kanuri, Luo, Nande, Somali, Twi, Wolof, Yoruba) generated using large language models (GPT-4o, GPT-4.5, Claude 3.5 Sonnet, Claude 3.7 Sonnet). Each sentence has
This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-t
International audience
LughaGen is a curated multilingual corpus for four Kenyan and East African languages: Swahili (sw),