Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
YanSonZhuCui
Hôte:avatar
The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, which is scarce for low-resource languages. This paper presents GigaSpeech 2, a large-scale, multi-domain, multilingual speech recognition corpus. It is designed for low-resource languages and does not rely on paired speech and text data. GigaSpeech 2 comprises about 30,000 hours of automatically transcribed speech, including Thai, Indonesian, and Vietnamese, gathered from unlabeled YouTube videos. We also introduce an automated pipeline for data crawling, transcription, and label refinement. Specifically, this pipeline involves Whisper for initial transcription, MMS for forced alignment, and multi-dimensional filtering for data quality assurance. A modified Noisy Student Training is developed to further refine flawed pseudo labels iteratively, thereby enhancing model performance. Experimental results on our manually transcribed evaluation set and two public test sets from Common Voice and FLEURS confirm our corpus's high quality and broad applicability. Notably, ASR models trained on GigaSpeech 2 can reduce the word error rate for Thai, Indonesian, and Vietnamese on our challenging and realistic YouTube test set by 25% to 40% compared to Whisper large-v3, with merely 10% model parameters. Furthermore, our ASR models trained on GigaSpeech 2 yield superior performance compared to commercial services. We hope that our newly introduced corpus and pipeline will open a new avenue for low-resource speech recognition and significantly facilitate research in this area. Accepted in ACL 2025 (Main)

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingComputation and LanguageSound

Similaires

The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African LanguagesFocused Crawling for Automated IsiXhosa Corpus BuildingOpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource LanguagesEmpowering Low-Resource Language ASR via Large-Scale Pseudo LabelingBreaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork LanguagesAfriMMT-EA: Multi-domain Machine Translation for Low-Resource East African Languages

The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages

We present the Thiomi Dataset, a large-scale multimodal corpus spanning ten African languages across

Focused Crawling for Automated IsiXhosa Corpus Building

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially

Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling

In this study, we tackle the challenge of limited labeled data for low-resource languages in ASR, fo

Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages

Automatic Speech Recognition (ASR) has reached impressive accuracy for high-resource languages, yet

AfriMMT-EA: Multi-domain Machine Translation for Low-Resource East African Languages