Logo Lanfrica

African News Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
David Ifeoluwa AdelaniJesujoba O. Alabi
Publisher:
Zenodo
Host:avatar

This consist of a monolingual news corpus for 19 languages from various sources like VOA, BBC, isolezwe etc. 

- The BBC corpus (except for Yoruba) was extracted from the castorini/afriberta-corpus, please cite the Small Data? No Problem! Exp… if you use it. A big thank you to Kelechi Ogueji for providing this corpus

- The VOA corpus was extracted from the MOT corpus, please cite the MOT paper if you use it. 

- The Isolezwe (xho, zul) was crawled as part of the Lacuna NER/POS project with Masakhane, please cite the MAFAND paper for that. 

- The nya data was part of the AI4D -- African Language Pr… 

- We thank Jonathan Mukiibi for providing the lug news corpus. 

- If you use the corpus for amh, hau, ibo, kin, lug, luo, pcm, swa, wol, yor, please cite our Multilingual language model…. We provide a description of the sources in the paper.