Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Igbo-TTS-Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
This dataset comprises audio recordings of Igbo speech aligned with textual transcriptions. The dataset is structured into 17 folders, each containing audio files and a corresponding audio-text mapping file. The audio clips are short, typically ranging from 2 to 45 seconds, and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file. The textual content used in this dataset originates from a variety of written and spoken sources in Igbo, including oral tradition narratives, cultural commentary, news reporting and current affairs, social discourse, and everyday speech samples. These texts were segmented into short utterances suitable for read speech and TTS modelling.

Visit

mozilladatacollective.com

Tasks

speech processingtext to speech

Languages

Igbo

Tags

mdcmozilla data collectiveTTSMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Igbo Speech Datasetigbo translation datasetIgbo Monolingual DatasetIgbo NER datasetHausa TTS DatasetKiswahili TTS Dataset

Igbo Speech Dataset

The most comprehensive Igbo speech dataset on HuggingFace - natural, real-world Igbo from native spe

igbo translation dataset

This dataset card aims to be a base template for new datasets. It has been generated using this raw

Igbo Monolingual Dataset

A dataset is a collection of Monolingual Igbo sentences.

Igbo NER dataset

Igbo Named Entity Recognition Dataset

Hausa TTS Dataset

This dataset contains Hausa language text-to-speech (TTS) recordings from multiple speakers. It incl

Kiswahili TTS Dataset

The dataset contains Kiswahili text and audio files. The dataset contains 7,108 text files and audio files. The Kiswahili dataset was created from an open-source non-copyrighted material: Kiswahili audio Bible. The authors permit use for non-profit, educational, a