Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Gender-Balanced and Region-Diverse Luganda Speech Dataset for Automatic Speech Recognition

Domaine:

natural language processing

Type de record:

dataset
Créateur:
KagKanNakatumba-Nabende, JoyceNab
Éditeur:
Mak
Éditeur:
Men
Hôte:avatar
This dataset contains a gender, age, and region-balanced Luganda speech corpus designed for bias-aware and fair Automatic Speech Recognition (ASR) research. The dataset includes speech recordings collected from diverse speakers across the Central, Eastern, Northern, and Western regions of Uganda. Each audio file is accompanied by demographic metadata including speaker gender, age group, and region of origin. The data has been curated and filtered to ensure equal representation of male and female speakers within every region–age group. The dataset is intended to support research on fairness, bias mitigation, and performance evaluation in Luganda ASR systems, as well as linguistic and sociophonetic analysis of regional and demographic variation in Luganda speech.

Visit

doi.orgdata.mendeley.com

Tasks

automatic speech recognitionspeech processing

Languages

Ganda

Tags

Artificial IntelligenceNatural Language ProcessingSpeech Recognition

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode