Logo Lanfrica

pikeloinc/african-francophone-speech

Domaine:

natural language processing

Type de record:

dataset
Créateur:
pik
Hôte:
A dataset of speech patterns, pronunciation errors, and linguistic mistakes made by African Francophone speakers, designed to improve voice AI and speech recognition for African accents. # African Francophone Speech A dataset of speech patterns, pronunciation errors, and linguistic mistakes made by African Francophone speakers. The goal is to improve voice AI and automatic speech recognition (ASR) for African French accents. Mainstream speech models are trained mostly on European French. African Francophone varieties remain underrepresented. This project collects recordings, transcripts, and structured error annotations so models can better understand African French speakers. ## Repository layout | Path | Contents | | --- | --- | | `data/raw/` | Original audio and transcripts | | `data/processed/` | Normalized audio and aligned transcripts | | `metadata/` | Speaker and recording metadata | | `annotations/` | Error labels (pronunciation, grammar, vocabulary, fluency, accent) | | `lexicons/` | Lexical resources and common mispronunciations | | `benchmarks/` | Evaluation sets and reported metrics | | `scripts/` | Preparation and validation utilities | | `docs/` | Annotation guidelines and project documentation | See DATASET.md for formats and schema. ## Contributing Read CONTRIBUTING.md and the annotation guidelines before submitting data or labels. ## License This repository is released under the MIT License.