Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Pretoria Sepedi Corpus (Gold Standard)

Type de record:

dataset
Éditeur:
Department of African Languages - University of Pretoria
Hôte:avatar
A section of the Pretoria Sepedi Corpus for POS, manually checked for POS tags.

Visit

hdl.handle.net

Tasks

part of speech tagging

Languages

Sotho, Northern

Similaires

Pretoria Sepedi Corpus POS taggedBBC Igbo-Pidgin Gold-Standard NLP Corpus<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>Pretoria Tshivenda CorpusSepedi NER CorpusLwazi Sepedi TTS corpus

Pretoria Sepedi Corpus POS tagged

The tagged Pretoria Sepedi Corpus for part-of-speech (POS) tagging. For grammtical anlysis morpholog

BBC Igbo-Pidgin Gold-Standard NLP Corpus

This corpus contains high-quality BBC Igbo and BBC Pidgin news snippets annotated by Bytte AI for intent, sentiment, content quality, sentence segmentation, and Named Entity Recognition (NER). Comprising 63 Igbo and 91 Pidgin samples per task, it provides a rich

<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>

This corpus is a high-quality, manually annotated collection of BBC Igbo and BBC Pidgin

Pretoria Tshivenda Corpus

Collection of texts for general linguistic research, in particular for lexicography

Sepedi NER Corpus

The Sepedi Ner Corpus is a Sepedi dataset developed by The Centre for Text Technology (CTexT), North

Lwazi Sepedi TTS corpus

Orthographic and phonemically aligned transcriptions