Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NCHLT Siswati Auxiliary Speech Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Febe de WetLaura MartinusJaco Badenhorst
Editor:
Charl van HeerderEtienne BarnardMarelie DavelAlta de Waal
Publisher:
CSIR Meraka InstituteNorth-West University
Host:avatar
The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven official languages. Transcriptions are provided in XML format.

Visit

hdl.handle.net

Tasks

speech processing

Languages

Swati

Tags

Siswati; Speech corpora; Transcribed

Licenses

Creative Commons Attribution 3.0 Unported (CC BY 3.0): https://creativecommons.org/licenses/by/3.0/legalcode

Similar

NCHLT Speech Corpus -- siSwatiNCHLT Siswati Speech CorpusNCHLT Speech Corpus -- siSwatiNCHLT Sepedi Auxiliary Speech CorpusNCHLT Setswana Auxiliary Speech CorpusNCHLT isiZulu Auxiliary Speech Corpus

NCHLT Speech Corpus -- siSwati

This is the siSwati language part of the NCHLT Speech Corpus of the South African languages. Languag

NCHLT Siswati Speech Corpus

Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test sui

NCHLT Speech Corpus -- siSwati

This is the siSwati language part of the NCHLT Speech Corpus of the South African languages. Languag

NCHLT Sepedi Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT Setswana Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT isiZulu Auxiliary Speech Corpus

This is the Zulu language split of the NCHLT speech corpus (nchlt-clean split). It containes 56 hour