Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Sepedi NER Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
com
Hôte:
The Sepedi Ner Corpus is a Sepedi dataset developed by The Centre for Text Technology (CTexT), North-West University, South Africa. The data is based on documents from the South African goverment domain and crawled from gov.za websites. It was created to support NER task for Sepedi language. The dataset uses CoNLL shared task annotation standards. Supported Tasks and Leaderboards

Visit

huggingface.co

Tasks

information extractionnamed entity recognition

Languages

Sotho, Northern

Licenses

other

Similaires

Sepedi NerNCHLT Sepedi Speech CorpusNCHLT Speech Corpus -- SepediLwazi Sepedi ASR corpusLwazi Sepedi TTS corpusNCHLT Speech Corpus -- Sepedi

Sepedi Ner

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

NCHLT Sepedi Speech Corpus

Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test sui

NCHLT Speech Corpus -- Sepedi

This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages. Language

Lwazi Sepedi ASR corpus

Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems.

Lwazi Sepedi TTS corpus

Orthographic and phonemically aligned transcriptions

NCHLT Speech Corpus -- Sepedi

This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages. Language