Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

High quality TTS data for four South African languages (af, st, tn, xh)

Domain:

natural language processing

Record type:

dataset
Editor:
GoogleNorth-West University
Publisher:
GoogleNorth-West University
Host:avatar
This data set contains multi-speaker TTS high quality transcribed audio data for four languages of South Africa: Afrikaans, Sesotho, Setswana and isiXhosa. The data set consists of wave files, and a TSV file transcribing the audio. In each folder, the file line_index.tsv contains a FileID, which in turn contains the UserID and the Transcription of audio in the file. The data set has had some quality checks, but there might still be errors. This data set was collected by as a collaboration between North-West University and Google. See LICENSE.txt file for license information. Copyright 2017 Google, Inc.

Visit

hdl.handle.net

Tasks

speech processingtext to speech

Languages

AfrikaansSetswanaSotho, SouthernXhosa

Tags

TTS

Licenses

Attribution-ShareAlike 4.0 International (CC BY-SA 4.0): https://creativecommons.org/licenses/by-sa/4.0/

Similar

SLR32 – High Quality TTS Data for Four South African Languages

SLR32 – High Quality TTS Data for Four South African Languages

Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/ This dataset contains mult