Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks

Domain:

natural language processing

Record type:

paperdataset
Creator:
Ma,KoiKarZen
Host:avatar
This paper introduces FLEURS-R, a speech restoration applied version of the Few-shot Learning Evaluation of Universal Representations of Speech (FLEURS) corpus. FLEURS-R maintains an N-way parallel speech corpus in 102 languages as FLEURS, with improved audio quality and fidelity by applying the speech restoration model Miipher. The aim of FLEURS-R is to advance speech technology in more languages and catalyze research including text-to-speech (TTS) and other speech generation tasks in low-resource languages. Comprehensive evaluations with the restored speech and TTS baseline models trained from the new corpus show that the new corpus obtained significantly improved speech quality while maintaining the semantic contents of the speech. The corpus is publicly released via Hugging Face.

Visit

arxiv.org

Tasks

speech processingtext to speech

Tags

Computation and LanguageArtificial IntelligenceSoundAudio and Speech Processing

Similar

GPTs Are Multilingual Annotators for Sequence Generation TasksCS-FLEURS: A Massively Multilingual and Code-Switched Speech DatasetZambezi Voice: A Multilingual Speech Corpus for Zambian LanguagesMultilingual Intermediate Tasks for Zero-Shot Cross-Lingual Transfer in XTREME-RCommon Voice: A Massively-Multilingual Speech CorpusNCHLT Auxiliary Speech Corpus - Multilingual

GPTs Are Multilingual Annotators for Sequence Generation Tasks

Data annotation is an essential step for constructing new datasets. However, the conventional approa

CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset

We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition a

Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages

This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages. It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours) and labelled data (over 80 hours) consistin

Multilingual Intermediate Tasks for Zero-Shot Cross-Lingual Transfer in XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Common Voice: A Massively-Multilingual Speech Corpus

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identi

NCHLT Auxiliary Speech Corpus - Multilingual

This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data S