Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Monolingual isiXhosa corpus

Domain:

natural language processing

Record type:

dataset
Creator:
McKellar, Cindy
Publisher:
North-West University - Centre for Text Technology (CTexT)
Host:avatar
Monolingual corpus for isiXhosa. The data is given as a single UTF-8 text file, with each segment on a newline. The dataset contains existing data sourced for the DAC funded Autshumato project as well as new data sourced for the SADiLaR: Parallel corpora for English into isiXhosa project. The data comprises a total of 233,192 segments with 2,424,706 isiXhosa words. The first 48,213 lines of the data file contain the original Autshumato project data and the remaining part of the file contains the new data.

Visit

hdl.handle.net

Tasks

language modeling

Languages

Xhosa

Tags

Monolingual corpus; isiXhosa

Licenses

Creative Commons Attribution 4.0 International: http://creativecommons.org/licenses/by/4.0/

Similar

Autshumato Monolingual isiXhosa Monolingual corpusMonolingual Siswati CorpusAcoli Monolingual CorpusLumasaba Monolingual CorpusKiswahili Monolingual CorpusLuganda Monolingual Corpus

Autshumato Monolingual isiXhosa Monolingual corpus

Monolingual corpus for isiXhosa. The data is given as a single UTF-8 text file, with each segment on

Monolingual Siswati Corpus

Monolingual corpus for SiSwati. The data is given as a single UTF-8 text file, with each segment on

Acoli Monolingual Corpus

Acoli is a very low-resourced language spoken in parts of Northern Uganda. This dataset contains 40,

Lumasaba Monolingual Corpus

Lumasaba sometimes known as Lugisu is a Bantu language spoken in the Eastern part of Uganda. This da

Kiswahili Monolingual Corpus

This dataset contains 100,000 Kiswahili sentences. For more information on how the dataset

Luganda Monolingual Corpus

This dataset contains 100,000 Luganda sentences. For more information on how the dataset wa