Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TWB Voice 1.0 - Hausa

Domaine:

natural language processing

Type de record:

dataset
Créateur:
CLEAR Global
Hôte:
TWB Voice 1.0 - Hausa is the Hausa language portion of the TWB Voice 1.0 multilingual speech corpus, created by CLEAR Global (formerly Translators without Borders). It contains approximately 58 hours of read speech recorded by native Hausa speakers through the TWB Voice platform. The dataset includes 36,665 recordings across train, dev, test, rejected, and pending splits, with transcriptions and speaker demographic metadata (age, gender, education level, country of origin). Audio is in WAV format at 48kHz. The dataset was created to support automatic speech recognition (ASR) development for underrepresented languages, with funding from the Patrick J. McGovern Foundation.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processing

Languages

Hausa

Tags

mdcmozilla data collectiveASRWAVTSV

Licenses

Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)

Similaires

CLEAR-Global/TWB-Voice-Hausa-TTS-1.0CLEAR-Global/TWB-voice-TTS-Hausa-1.0-samplesetTWB Voice Hausa TTS Dataset 1.0 - Sample SetTWB Voice 1.0 - KanuriCLEAR-Global/TWB-Voice-1.0TWB Voice 1.0 - Shuwa Arabic

CLEAR-Global/TWB-Voice-Hausa-TTS-1.0

CLEAR-Global/TWB-voice-TTS-Hausa-1.0-sampleset

Dataset Summary

TWB Voice Hausa TTS Dataset 1.0 - Sample Set

TWB Voice Hausa TTS 1.0 Sample Set is a high-quality text-to-speech corpus containing read speech data in Hausa, recorded by a single female speaker under acoustically optimal conditions. This dataset represents 10% of the complete Hausa TTS dataset collected as

TWB Voice 1.0 - Kanuri

TWB Voice 1.0 - Kanuri is the Kanuri language portion of the TWB Voice 1.0 multilingual speech corpu

CLEAR-Global/TWB-Voice-1.0

TWB Voice 1.0 is a multilingual speech corpus containing read speech data in three languages from Ni

TWB Voice 1.0 - Shuwa Arabic

TWB Voice 1.0 - Shuwa Arabic is the Shuwa Arabic language portion of the TWB Voice 1.0 multilingual