Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NCHLT Afrikaans Auxiliary Speech Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Febe de WetLaura MartinusJaco Badenhorst
Editor:
Charl van HeerderEtienne BarnardMarelie DavelAlta de Waal
Publisher:
CSIR Meraka InstituteNorth-West University
Host:avatar
The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven official languages. Transcriptions are provided in XML format.

Visit

hdl.handle.net

Tasks

speech processing

Languages

Afrikaans

Tags

Afrikaans; Speech corpora; Transcribed

Licenses

Creative Commons Attribution 3.0 Unported (CC BY 3.0): https://creativecommons.org/licenses/by/3.0/legalcode

Similar

NCHLT isiZulu Auxiliary Speech CorpusNCHLT Auxiliary Speech Corpus - MultilingualNCHLT Xitsonga Auxiliary Speech CorpusNCHLT Setswana Auxiliary Speech CorpusNCHLT Sepedi Auxiliary Speech CorpusNCHLT Tshivenda Auxiliary Speech Corpus

NCHLT isiZulu Auxiliary Speech Corpus

This is the Zulu language split of the NCHLT speech corpus (nchlt-clean split). It containes 56 hour

NCHLT Auxiliary Speech Corpus - Multilingual

This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data S

NCHLT Xitsonga Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT Setswana Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT Sepedi Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT Tshivenda Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o