Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexi
Aligned parallel corpora for the language pair English-Sepedi. The data is given as two separate UTF
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Te
This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages. Language