Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NCHLT Tshivenda Text Corpora

Domain:

natural language processing

Record type:

dataset
Creator:
Martin PuttkammerMartin SchlemmerWikus PienaarRuan Bekker
Publisher:
North-West UniversityCentre for Text Technology (CTexT)
Host:avatar
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project.

Visit

hdl.handle.net

Languages

Venda

Licenses

Creative Commons Attribution 2.5 South Africa License: http://creativecommons.org/licenses/by/2.5/za/legalcode

Similar

NCHLT Tshivenda Annotated Text CorporaNCHLT English Text CorporaNCHLT Sepedi Text CorporaNCHLT Setswana Text CorporaNCHLT Sesotho Text CorporaNCHLT Afrikaans Text Corpora

NCHLT Tshivenda Annotated Text Corpora

Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Te

NCHLT English Text Corpora

Collection consisting of a clean corpus, lexicon, frequency list and named-entity lists developed d

NCHLT Sepedi Text Corpora

Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexi

NCHLT Setswana Text Corpora

Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexi

NCHLT Sesotho Text Corpora

Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexi

NCHLT Afrikaans Text Corpora

Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexi