Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

isiXhosa-TTS-Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
This dataset comprises audio recordings of isiXhosa speech aligned with textual transcriptions. The dataset is structured into 24 folders, each containing audio files and a corresponding audio-text mapping file. The audio clips are short, typically ranging from 1 to 13 seconds, and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file. The textual content used in this dataset originates from written isiXhosa sources published on the indigenous-language blogging platform IndigenousBlogs (indigenousblogs.com), which hosts original content authored by isiXhosa-speaking bloggers across a range of topics, including narrative texts, opinion pieces, cultural commentary, and everyday informational content. These texts were segmented into short utterances suitable for read speech and TTS modelling.

Visit

mozilladatacollective.com

Tasks

speech processingtext to speech

Languages

Xhosa

Tags

mdcmozilla data collectiveTTSMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Lwazi isiXhosa TTS corpusIsiXhosa multi-speaker TTS corpusLwazi III isiXhosa TTS CorpusLwazi II isiXhosa TTS CorpusisiXhosa NLP DatasetisiXhosa ASR Dataset

Lwazi isiXhosa TTS corpus

Orthographic and phonemically aligned transcriptions

IsiXhosa multi-speaker TTS corpus

The aim of this corpus was to investigate the implementation of a high-quality TTS system using mult

Lwazi III isiXhosa TTS Corpus

Complete audio recordings with orthographic transcriptions. TTS corpus for standard SA dialect. This

Lwazi II isiXhosa TTS Corpus

Orthographic and phonemically aligned transcriptions.

isiXhosa NLP Dataset

A high-quality, comprehensive isiXhosa (Xhosa) NLP training dataset carefully collected, cleaned, an

isiXhosa ASR Dataset

This dataset contains speech recordings and transcriptions for isiXhosa, one of South Africa's offic