Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Discrete vs Continuous Audio Token Representations in Cross-Lingual Transfer Accuracy on CommonVoice Low-Resource Benchmark

Domain:

natural language processing
Creator:
SOV
Publisher:
Zenodo
Host:avatar
This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages. We build on wav2vec 2.0 which is trained by solving a contrastive task over masked latent speech representations and jointly learns a quantization of the latents shared across languages. The resulting model is fine-tuned on labeled data and experiments show that cross-lingual pretraining significantly outperforms monolingual pretraining. On the CommonVoice benchmark, XLSR shows a relative phoneme error rate reduction of 72\% compared to Research goal: How do discrete audio token representations compare to continuous features in cross-lingual transfer accuracy on the CommonVoice low-resource benchmark? Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 7.8/10. This report was generated autonomously by SOVEREIGN Research Kernel, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.8/10.

Visit

doi.orgzenodo.org

Tasks

automatic speech recognitionspeech processingtransfer learning

Tags

discreteaudiotokenrepresentationscontinuousfeaturescross-lingualtransfer

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode