Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NCHLT Sepedi Speech Corpus

Record type:

dataset
Creator:
Charl van HeerdenEtienne BarnardJaco BadenhorstMarelie Davel
Editor:
Willem BassonNic de VriesFebe de WetThipe Modipa
Publisher:
Meraka Institute, CSIRNorth-West University
Host:avatar
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers.

Visit

hdl.handle.net

Tasks

speech processing

Languages

Sotho, Northern

Licenses

Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode

Similar

NCHLT Speech Corpus -- SepediNCHLT Speech Corpus -- SepediNCHLT Sepedi Auxiliary Speech CorpusNCHLT Sepedi Phrase Chunk Annotated CorpusNCHLT Sepedi Named Entity Annotated CorpusNCHLT speech corpus

NCHLT Speech Corpus -- Sepedi

This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages. Language

NCHLT Speech Corpus -- Sepedi

This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages. Language

NCHLT Sepedi Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT Sepedi Phrase Chunk Annotated Corpus

Phrase chunk annotated data for the NCHLT Text Resource Development: Phase II Project. The phrase ch

NCHLT Sepedi Named Entity Annotated Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated wi

NCHLT speech corpus

The NCHLT speech corpus of the South African languages