Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TWB Voice 1.0 - Kanuri

Domain:

natural language processing

Record type:

dataset
Creator:
CLEAR Global
Host:
TWB Voice 1.0 - Kanuri is the Kanuri language portion of the TWB Voice 1.0 multilingual speech corpus, created by CLEAR Global (formerly Translators without Borders). It contains approximately 52 hours of read speech recorded by native Kanuri speakers through the TWB Voice platform. The dataset includes 27,723 recordings across train, dev, test, rejected, and pending splits, with transcriptions and speaker demographic metadata (age, gender, education level, country of origin). Audio is in WAV format at 48kHz. The dataset was created to support automatic speech recognition (ASR) development for underrepresented languages, with funding from the Patrick J. McGovern Foundation.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processing

Languages

KanembuKanuri, MangaKanuri, Yerwa

Tags

mdcmozilla data collectiveASRWAVTSV

Licenses

Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)

Similar

CLEAR-Global/TWB-Voice-Kanuri-TTS-1.0CLEAR-Global/TWB-voice-TTS-Kanuri-1.0-samplesetTWB Voice 1.0 - HausaCLEAR-Global/TWB-Voice-1.0TWB Voice 1.0 - Shuwa ArabicCLEAR-Global/TWB-Voice-Hausa-TTS-1.0

CLEAR-Global/TWB-Voice-Kanuri-TTS-1.0

CLEAR-Global/TWB-voice-TTS-Kanuri-1.0-sampleset

Dataset Summary

TWB Voice 1.0 - Hausa

TWB Voice 1.0 - Hausa is the Hausa language portion of the TWB Voice 1.0 multilingual speech corpus,

CLEAR-Global/TWB-Voice-1.0

TWB Voice 1.0 is a multilingual speech corpus containing read speech data in three languages from Ni

TWB Voice 1.0 - Shuwa Arabic

TWB Voice 1.0 - Shuwa Arabic is the Shuwa Arabic language portion of the TWB Voice 1.0 multilingual

CLEAR-Global/TWB-Voice-Hausa-TTS-1.0