Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Setswana Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Kok
Host:
I used my Extracting and Corpus Creation notebooks to create this corpus.

Visit

www.kaggle.com

Languages

Setswana

Similar

Setswana NER CorpusSetswana Ner CorpusAutshumato Monolingual Setswana CorpusLwazi Setswana TTS corpusNCHLT Setswana Speech CorpusSetswana Genre Classification Corpus

Setswana NER Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated wi

Setswana Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Autshumato Monolingual Setswana Corpus

Monolingual corpus for Setswana. The data is given as a single UTF-8 text file, with each segment on

Lwazi Setswana TTS corpus

Orthographic and phonemically aligned transcriptions

NCHLT Setswana Speech Corpus

Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test sui

Setswana Genre Classification Corpus

Contains training and testing data for Genre Classification for Setswana.