Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Lumasaba Monolingual Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
NabMuzBabirye, ClaireMukiibi, Jonathan
Editor:
Mukiibi, Jonathan
Publisher:
Har
Host:avatar
Lumasaba sometimes known as Lugisu is a Bantu language spoken in the Eastern part of Uganda. This dataset contains a total of 39,999 sentences. The sentences are split into two separate files. One file contains 20,764 sentences from the Northern dialect and another one contains 19,235 sentences from the Southern dialect. This dataset was compiled by a team of Linguists and researchers from the Makerere AI and Data Science Research Lab and Marconi Research and Innovation Lab at Makerere University. This dataset was created with support from Lacuna Fund.

Visit

doi.orgdataverse.harvard.edu

Languages

Masaaba

Tags

Arts and HumanitiesComputer and Information ScienceEngineeringNatural Language ProcessingMonolingualMasaba, Gisu

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Autshumato Monolingual isiXhosa Monolingual corpusMonolingual isiXhosa corpusLuganda Monolingual CorpusAcoli Monolingual CorpusMonolingual Siswati CorpusKiswahili Monolingual Corpus

Autshumato Monolingual isiXhosa Monolingual corpus

Monolingual corpus for isiXhosa. The data is given as a single UTF-8 text file, with each segment on

Monolingual isiXhosa corpus

Monolingual corpus for isiXhosa. The data is given as a single UTF-8 text file, with each segment on

Luganda Monolingual Corpus

This dataset contains 100,000 Luganda sentences. For more information on how the dataset wa

Acoli Monolingual Corpus

Acoli is a very low-resourced language spoken in parts of Northern Uganda. This dataset contains 40,

Monolingual Siswati Corpus

Monolingual corpus for SiSwati. The data is given as a single UTF-8 text file, with each segment on

Kiswahili Monolingual Corpus

This dataset contains 100,000 Kiswahili sentences. For more information on how the dataset