Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TWB Voice 1.0 - Hausa

Domain:

natural language processing

Record type:

dataset
Creator:
CLEAR Global
Host:
TWB Voice 1.0 - Hausa is the Hausa language portion of the TWB Voice 1.0 multilingual speech corpus, created by CLEAR Global (formerly Translators without Borders). It contains approximately 58 hours of read speech recorded by native Hausa speakers through the TWB Voice platform. The dataset includes 36,665 recordings across train, dev, test, rejected, and pending splits, with transcriptions and speaker demographic metadata (age, gender, education level, country of origin). Audio is in WAV format at 48kHz. The dataset was created to support automatic speech recognition (ASR) development for underrepresented languages, with funding from the Patrick J. McGovern Foundation.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processing

Languages

Hausa

Tags

mdcmozilla data collectiveASRWAVTSV

Licenses

Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)

Similar

CLEAR-Global/TWB-Voice-Hausa-TTS-1.0CLEAR-Global/TWB-voice-TTS-Hausa-1.0-samplesetTWB Voice Hausa TTS Dataset 1.0 - Sample SetTWB Voice 1.0 - KanuriTWB Voice 1.0 - Shuwa ArabicCLEAR-Global/TWB-Voice-1.0

CLEAR-Global/TWB-Voice-Hausa-TTS-1.0

CLEAR-Global/TWB-voice-TTS-Hausa-1.0-sampleset

Dataset Summary

TWB Voice Hausa TTS Dataset 1.0 - Sample Set

TWB Voice Hausa TTS 1.0 Sample Set is a high-quality text-to-speech corpus containing read speech data in Hausa, recorded by a single female speaker under acoustically optimal conditions. This dataset represents 10% of the complete Hausa TTS dataset collected as

TWB Voice 1.0 - Kanuri

TWB Voice 1.0 - Kanuri is the Kanuri language portion of the TWB Voice 1.0 multilingual speech corpu

TWB Voice 1.0 - Shuwa Arabic

TWB Voice 1.0 - Shuwa Arabic is the Shuwa Arabic language portion of the TWB Voice 1.0 multilingual

CLEAR-Global/TWB-Voice-1.0

TWB Voice 1.0 is a multilingual speech corpus containing read speech data in three languages from Ni