Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NCHLT Siswati Text Corpora

Domain:

natural language processing

Record type:

dataset
Creator:
Martin PuttkammerMartin SchlemmerWikus PienaarRuan Bekker
Publisher:
North-West UniversityCentre for Text Technology (CTexT)
Host:avatar
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project.

Visit

hdl.handle.net

Languages

Swati

Licenses

Creative Commons Attribution 2.5 South Africa License: http://creativecommons.org/licenses/by/2.5/za/legalcode

Similar

NCHLT Siswati Annotated Text CorporaNCHLT isiXhosa Text CorporaNCHLT Sepedi Text CorporaNCHLT Tshivenda Text CorporaNCHLT English Text CorporaNCHLT isiZulu Text Corpora

NCHLT Siswati Annotated Text Corpora

Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Te

NCHLT isiXhosa Text Corpora

Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexi

NCHLT Sepedi Text Corpora

Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexi

NCHLT Tshivenda Text Corpora

Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexi

NCHLT English Text Corpora

Collection consisting of a clean corpus, lexicon, frequency list and named-entity lists developed d

NCHLT isiZulu Text Corpora

Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexi