Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NCHLT Sesotho Auxiliary Speech Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Febe de WetLaura MartinusJaco Badenhorst
Editor:
Charl van HeerderEtienne BarnardMarelie DavelAlta de Waal
Publisher:
CSIR Meraka InstituteNorth-West University
Host:avatar
The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven official languages. Transcriptions are provided in XML format.

Visit

hdl.handle.net

Tasks

speech processing

Languages

Sotho, Southern

Tags

Sesotho; Speech corpora; Transcribed

Licenses

Creative Commons Attribution 3.0 Unported (CC BY 3.0): https://creativecommons.org/licenses/by/3.0/legalcode

Similar

NCHLT isiXhosa Auxiliary Speech CorpusNCHLT isiZulu Auxiliary Speech CorpusNCHLT Setswana Auxiliary Speech CorpusNCHLT Siswati Auxiliary Speech CorpusNCHLT Afrikaans Auxiliary Speech CorpusNCHLT Xitsonga Auxiliary Speech Corpus

NCHLT isiXhosa Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT isiZulu Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT Setswana Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT Siswati Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT Afrikaans Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o

NCHLT Xitsonga Auxiliary Speech Corpus

The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven o