Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An Overview

Domaine:

natural language processing

Type de record:

paper
This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and automatic speech recognition applications for languages and dialects of South and Southeast Asia, Africa, Europe and South America. The paper describes the methodology used for developing such corpora and presents some of our findings that could benefit under-represented language communities.

Visit

arxiv.org

Tasks

text to speechautomatic speech recognitionspeech processing

Languages

AfrikaansPidgin, NigerianSetswanaSotho, SouthernXhosa

Similaires

Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and DialectsGhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian LanguagesOpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource LanguagesLIMBA: An Open-Source Framework for the Preservation and Valorization of Low-Resource Languages using Generative ModelsBuilding Corpora for Low-Resource Kenyan LanguagesOpen but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages

Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects

DVoice is a community initiative that aims to provide African languages and dialects with d

GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages

Low resource languages present unique challenges for natural language processing due to the limited

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially

LIMBA: An Open-Source Framework for the Preservation and Valorization of Low-Resource Languages using Generative Models

Minority languages are vital to preserving cultural heritage, yet they face growing risks of extinct

Building Corpora for Low-Resource Kenyan Languages

Natural Language Processing is a crucial frontier in artificial intelligence, with broad application

Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages

Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are ra