Logo Lanfrica

Luganda Monolingual Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Mukiibi, JonathanBabirye, ClaireTusubira, JeremyBateesa, Tobius
Editor:
Ekibiina Ky'Olulimi OlugandaBuganda Land BoardGambuuze
Publisher:
Zenodo
Host:avatar

This dataset contains 100,000 Luganda sentences. For more information on how the dataset was created, please check out our paper published at AfricaNLP. We want to thank the Department of African Languages, Makerere University. We would like to thank the Buganda Kingdom for partnering with us and also for the support on the monolingual text collection for Luganda through its agencies. We would also like to acknowledge the various organisations and individuals from which we sourced this data; the Independent News, Jerum Agency and Bonamix Translation Services Ltd. This dataset was created with support from Lacuna Fund, an initiative cofounded by The Rockefeller Foundation, Google.org, and Canada’s International Development Research Centre; Deutsche Gesellschaft fur Internationale ¨ Zusammenarbeit (GIZ) and Mozilla.