Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Gender-Balanced and Region-Diverse Luganda Speech Dataset for Automatic Speech Recognition

Domain:

natural language processing

Record type:

dataset
Creator:
KagKanNakatumba-Nabende, JoyceNab
Editor:
Mak
Publisher:
Men
Host:avatar
This dataset contains a gender, age, and region-balanced Luganda speech corpus designed for bias-aware and fair Automatic Speech Recognition (ASR) research. The dataset includes speech recordings collected from diverse speakers across the Central, Eastern, Northern, and Western regions of Uganda. Each audio file is accompanied by demographic metadata including speaker gender, age group, and region of origin. The data has been curated and filtered to ensure equal representation of male and female speakers within every region–age group. The dataset is intended to support research on fairness, bias mitigation, and performance evaluation in Luganda ASR systems, as well as linguistic and sociophonetic analysis of regional and demographic variation in Luganda speech.

Visit

doi.orgdata.mendeley.com

Tasks

automatic speech recognitionspeech processing

Languages

Ganda

Tags

Artificial IntelligenceNatural Language ProcessingSpeech Recognition

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Mozilla Luganda Automatic Speech RecognitionThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionSomali Automatic Speech Recognition DatasetAfriVox: An African benchmark dataset for Automatic Speech Translation and Speech RecognitionKARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

Mozilla Luganda Automatic Speech Recognition

Can you create an Automatic Speech Recognition model for Luganda?
The objective of this challenge is to create an automatic speech recognition model on Luganda. You will train your models on the complete Luganda dataset provided by Mozilla Common Voice and

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda. The corp

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the Un

Somali Automatic Speech Recognition Dataset

This dataset contains audio recordings and corresponding transcriptions in Somali, designed for auto

AfriVox: An African benchmark dataset for Automatic Speech Translation and Speech Recognition

This project creates a benchmark dataset for evaluating Automatic Speech Translation and Speech reco

KARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

While large-scale pre-trained models have significantly advanced multilingual Automatic Speech Recog