Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Naija-TTS-Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
This dataset comprises audio recordings of Nigerian Pidgin English speech aligned with textual transcriptions. The dataset is structured into 16 folders, each containing audio files and a corresponding audio-text mapping file. The audio clips are short, typically ranging from 1 to 38 seconds, and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file. The textual content used in this dataset originates from a variety of written and spoken sources in Nigerian Pidgin English, including narrative texts, conversational exchanges, news-style content, and everyday speech samples. These texts were segmented into short utterances suitable for read speech and TTS modelling.

Visit

mozilladatacollective.com

Tasks

speech processingtext to speech

Languages

Ghanaian Pidgin EnglishPidgin, Nigerian

Tags

mdcmozilla data collectiveTTSMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

slickcityceo/naija-ttsOdiabackend099/odiadev-tts-naija-aiMbosi-TTS-DatasetRw Tts DatasetYaka-TTS-DatasetEton-TTS-Dataset

slickcityceo/naija-tts

Nigeria's first trainable text to speech, Naija TTS is a text-to-speech application that converts En

Odiabackend099/odiadev-tts-naija-ai

# Protect.NG CrossAI - Voice-First Emergency Platform > Nigeria's first AI-powered voice-led emerge

Mbosi-TTS-Dataset

The dataset consists of paired audio and text data on Mbosi (mdw), a language spoken in Congo. The a

Rw Tts Dataset

Kinyarwanda (rw) text-to-speech dataset. Studio-recorded read speech aligned with transcriptions, co

Yaka-TTS-Dataset

Paired audio and text data on Yaka (also known as West Teke), a language spoken in Congo. The audio

Eton-TTS-Dataset

Eton-TTS-Dataset is a single-speaker scripted speech dataset dedicated to the documentation and tech