Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
RahReh
Hôte:avatar
We present the Pashto Common Voice corpus -- the first large-scale, openly licensed speech resource for Pashto, a language with over 60 million native speakers largely absent from open speech technology. Through a community effort spanning 2022-2025, the corpus grew from 1.5 hours and 5 contributors to 147 total hours and 1,483 unique speakers across ten Mozilla Common Voice releases (CV14-CV23). Speaker participation increased approximately 108-fold between CV17 and CV18, coinciding with a VOA Pashto broadcast campaign. We describe the full methodology: interface localisation, Wikipedia-based sentence extraction with automated filtering, phonemically targeted contributions for the four most frequently dropped Pashto characters, and multi-channel community outreach. MCV23 contains 107,781 clips (60,337 validated; 82.33 validated hours) across 13 content domains. Fine-tuning Whisper Base on the MCV20 yields 13.4% WER on the MCV20 test split, against the published Whisper Base zero-shot WER of 99.0% on Pashto. Submitted to Interspeech 2026

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and Language

Similaires

Common Voice: A Massively-Multilingual Speech CorpusA Very Low Resource Language Speech Corpus for Computational Language Documentation ExperimentsTowards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource LanguageRIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language ProcessingDynAg Open Voice Dataset for Low-Resource Bihari LanguagesGiving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus

Common Voice: A Massively-Multilingual Speech Corpus

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identi

A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments

Most speech and language technologies are trained with massive amounts of speech and text informatio

Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language

The increase in technological adoption worldwide comes with demands for novel tools to be used by th

RIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

This dataset consists of a curated collection of high-fidelity, field-recorded audio samples develop

DynAg Open Voice Dataset for Low-Resource Bihari Languages

jjkhkjhkjh

Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus

We present a unified pipeline for synthesizing high-quality Quechua and Spanish speech for the Peruv