Building Inclusive AI for African Languages: Exploring the Mozilla Data Collective Hausa Speech Dataset
Licence MIT
Dataset Structure
Open Hausa Language AI Dataset (OHLAD)
│
├── audio/
│ ├── speaker_001/
│ ├── speaker_002/
│ └── ...
├── transcripts/
├── metadata/
├── consent_forms/
├── documentation/
└── LICENSE1. Project Proposal
Writing
Open Hausa Language AI Dataset (OHLAD)
Project Overview
The Open Hausa Language AI Dataset (OHLAD) is a community-driven initiative led by the Africa Cybersecurity and Criminals Investigation OSINT Lab (ACOL) to develop a high-quality, ethically collected Hausa language dataset for artificial intelligence research and development.
Objectives
Develop an open Hausa speech and text dataset.
Support AI applications including speech recognition, text-to-speech, machine translation, OCR, and conversational AI.
Promote digital inclusion for Hausa speakers.
Encourage community participation in African language technology.
Scope
The project will collect:
Speech recordings from native Hausa speakers.
Accurate Hausa transcripts.
Metadata describing recordings and speakers (without personally identifying contributors).
Documentation for researchers and developers.
Ethical Principles
Participation is voluntary.
Contributors provide informed consent.
Personal information is protected.
Data is collected and shared according to the selected project license.
Expected Impact
The dataset will help researchers, educators, startups, and open-source communities build better AI tools for the Hausa language and contribute to multilingual AI development across Africa.
2. Volunteer Consent Form
Writing
Volunteer Consent Form
I voluntarily agree to participate in the Open Hausa Language AI Dataset (OHLAD) project.
I understand that:
My voice recordings and transcripts will be used for AI research and language technology.
My participation is voluntary.
I may stop participating at any time before my data is incorporated into a released dataset, subject to the project's data management policy.
My name does not need to appear in the published dataset unless …