Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NCHLT Tshivenda Auxiliary Speech Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Febe de WetLaura MartinusJaco Badenhorst
Editor:
Charl van HeerderEtienne BarnardMarelie DavelAlta de Waal
Publisher:
CSIR Meraka InstituteNorth-West University
Host:avatar
The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven official languages. Transcriptions are provided in XML format.

Visit

hdl.handle.net

Tasks

speech processing

Languages

Venda

Tags

Tshivenda; Speech corpora; Transcribed

Licenses

Creative Commons Attribution 3.0 Unported (CC BY 3.0): https://creativecommons.org/licenses/by/3.0/legalcode

Similar

NCHLT Speech Corpus -- TshivendaNCHLT Speech Corpus -- TshivendaNCHLT Tshivenda Speech CorpusNCHLT isiZulu Auxiliary Speech CorpusNCHLT Auxiliary Speech Corpus - MultilingualNCHLT Afrikaans Auxiliary Speech Corpus

NCHLT Speech Corpus -- Tshivenda

This is the Tshivenda language part of the NCHLT Speech Corpus of the South African languages. Langu

NCHLT Speech Corpus -- Tshivenda

This is the Tshivenda language part of the NCHLT Speech Corpus of the South African languages. Langu

NCHLT Tshivenda Speech Corpus

Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test sui

NCHLT isiZulu Auxiliary Speech Corpus

This is the Zulu language split of the NCHLT speech corpus (nchlt-clean split). It containes 56 hour

NCHLT Auxiliary Speech Corpus - Multilingual

This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data S

NCHLT Afrikaans Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o