Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Discrete vs Continuous Audio Token Representations in Cross-Lingual Transfer Accuracy on CommonVoice Low-Resource Benchmark

Domaine:

natural language processing
Créateur:
SOV
Éditeur:
Zenodo
Hôte:avatar
This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages. We build on wav2vec 2.0 which is trained by solving a contrastive task over masked latent speech representations and jointly learns a quantization of the latents shared across languages. The resulting model is fine-tuned on labeled data and experiments show that cross-lingual pretraining significantly outperforms monolingual pretraining. On the CommonVoice benchmark, XLSR shows a relative phoneme error rate reduction of 72\% compared to Research goal: How do discrete audio token representations compare to continuous features in cross-lingual transfer accuracy on the CommonVoice low-resource benchmark? Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 7.8/10. This report was generated autonomously by SOVEREIGN Research Kernel, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.8/10.

Visit

doi.orgzenodo.org

Tasks

automatic speech recognitionspeech processingtransfer learning

Tags

discreteaudiotokenrepresentationscontinuousfeaturescross-lingualtransfer

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Zero-shot cross-lingual transfer accuracy in low-resource languages: multilingual vs. English-only intermediate tasksZero-shot Cross-lingual Transfer Accuracy in Multilingual vs. English Intermediate-Task Training for Low-Resource LanguagesDiscrete Audio Tokens Enhance Cross-Lingual Speech Recognition in Low-Resource LanguagesCross-lingual Transfer Accuracy and Task Similarity in Low-Resource LanguagesCross-lingual Transfer Accuracy in Low-Resource Languages via Task SimilarityDifficulty Level Impact on Zero-Shot Cross-Lingual Transfer Accuracy in XTREME Benchmark

Zero-shot cross-lingual transfer accuracy in low-resource languages: multilingual vs. English-only intermediate tasks

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Zero-shot Cross-lingual Transfer Accuracy in Multilingual vs. English Intermediate-Task Training for Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Discrete Audio Tokens Enhance Cross-Lingual Speech Recognition in Low-Resource Languages

This report synthesises findings from 13 peer-reviewed papers addressing the following research ques

Cross-lingual Transfer Accuracy and Task Similarity in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Cross-lingual Transfer Accuracy in Low-Resource Languages via Task Similarity

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Difficulty Level Impact on Zero-Shot Cross-Lingual Transfer Accuracy in XTREME Benchmark

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni