Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Pretoria Sepedi Corpus (Gold Standard)

Record type:

dataset
Publisher:
Department of African Languages - University of Pretoria
Host:avatar
A section of the Pretoria Sepedi Corpus for POS, manually checked for POS tags.

Visit

hdl.handle.net

Tasks

part of speech tagging

Languages

Sotho, Northern

Similar

Pretoria Sepedi Corpus POS taggedBBC Igbo-Pidgin Gold-Standard NLP Corpus<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>Pretoria Tshivenda CorpusSepedi NER CorpusNCHLT Speech Corpus -- Sepedi

Pretoria Sepedi Corpus POS tagged

The tagged Pretoria Sepedi Corpus for part-of-speech (POS) tagging. For grammtical anlysis morpholog

BBC Igbo-Pidgin Gold-Standard NLP Corpus

This corpus contains high-quality BBC Igbo and BBC Pidgin news snippets annotated by Bytte AI for intent, sentiment, content quality, sentence segmentation, and Named Entity Recognition (NER). Comprising 63 Igbo and 91 Pidgin samples per task, it provides a rich

<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>

This corpus is a high-quality, manually annotated collection of BBC Igbo and BBC Pidgin

Pretoria Tshivenda Corpus

Collection of texts for general linguistic research, in particular for lexicography

Sepedi NER Corpus

The Sepedi Ner Corpus is a Sepedi dataset developed by The Centre for Text Technology (CTexT), North

NCHLT Speech Corpus -- Sepedi

This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages. Language