Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Speech-to-text model comparison using XLS-R, XLSR-53, and Wav2Vec 2.0

Domaine:

natural language processing

Type de record:

paper
Créateur:
JuaAma
Éditeur:
Institute of Advanced Engineering and Science
Hôte:
Automatic speech recognition (ASR) systems have achieved significant progress in recent years; however, their performance remains limited for low-resource languages such as Indonesian. Multilingual ASR models are often expected to generalize across languages, yet they frequently underperform when applied to underrepresented languages without sufficient adaptation. This study presents a comparative evaluation of three ASR models—Wav2Vec 2.0, XLS-R, and XLSR-53—on Indonesian speech to analyze the impact of monolingual fine-tuning versus multilingual pretraining. The evaluation was conducted using approximately 28 hours of validated Indonesian speech from the Common Voice Corpus version 13. Model performance was assessed using word error rate (WER) without employing any external language model to ensure a fair comparison. Experimental results demonstrate that Wav2Vec 2.0, which is fine-tuned specifically for Indonesian, achieves substantially lower WER compared to the multilingual models. Qualitative analysis further confirms that multilingual models exhibit higher omission and substitution errors. These findings indicate that language-specific fine-tuning plays a more critical role than multilingual generalization in achieving accurate ASR for Indonesian. The results provide practical guidance for deploying ASR systems in low-resource language scenarios and highlight the importance of targeted model adaptation.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Licenses

https://creativecommons.org/licenses/by-sa/4.0

Similaires

Fon Automatic Speech Recognition Model (Wav2Vec2-Large-XLSR-53-Fon)marka/wav2vec-large-xls-r-300-ha-colab_2PaschalK/wav2vec-XLSR-swahiliad019el/Tamasheq-XLSR-53-1unza/xls-r-300m-nyanja-modelkingabzpro/wav2vec2-large-xlsr-53-wolof

Fon Automatic Speech Recognition Model (Wav2Vec2-Large-XLSR-53-Fon)

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Fon (or Fongbe) using the Fon Dataset. When using this model, make sure that your speech input is sampled at 16kHz.

marka/wav2vec-large-xls-r-300-ha-colab_2

PaschalK/wav2vec-XLSR-swahili

ad019el/Tamasheq-XLSR-53-1

unza/xls-r-300m-nyanja-model

kingabzpro/wav2vec2-large-xlsr-53-wolof