Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
BanAkkDevNar
Hôte:avatar
Automatic speech recognition for low-resource languages remains fundamentally constrained by the scarcity of labeled data and computational resources required by state-of-the-art models. We present a systematic investigation into cross-lingual continuous pretraining for low-resource languages, using Perso-Arabic languages (Persian, Arabic, and Urdu) as our primary case study. Our approach demonstrates that strategic utilization of unlabeled speech data can effectively bridge the resource gap without sacrificing recognition accuracy. We construct a 3,000-hour multilingual corpus through a scalable unlabeled data collection pipeline and employ targeted continual pretraining combined with morphologically-aware tokenization to develop a 300M parameter model that achieves performance comparable to systems 5 times larger. Our model outperforms Whisper Large v3 (1.5B parameters) on Persian and achieves competitive results on Arabic and Urdu despite using significantly fewer parameters and substantially less labeled data. These findings challenge the prevailing assumption that ASR quality scales primarily with model size, revealing instead that data relevance and strategic pretraining are more critical factors for low-resource scenarios. This work provides a practical pathway toward inclusive speech technology, enabling effective ASR for underrepresented languages without dependence on massive computational infrastructure or proprietary datasets. Accepted in AACL IJCNLP 2025

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageAudio and Speech Processing

Similaires

Scaling Cross-Lingual NER Performance with Unlabeled Target Data in Low-Resource LanguagesCross-lingual NER Transfer Accuracy and Unlabeled Target-Language Data Volume in Low-Resource LanguagesCross-lingual NER Performance with Unlabeled Target Data in Low-Resource SettingsScaling Out-of-Domain Unlabeled Data for Robust Cross-Lingual NER Alignment in Low-Resource SettingsCross-Lingual NER for Financial Transaction Data in Low-Resource LanguagesIsomorphic Cross-lingual Embeddings for Low-Resource Languages

Scaling Cross-Lingual NER Performance with Unlabeled Target Data in Low-Resource Languages

To better tackle the named entity recognition (NER) problem on languages with little/no labeled data

Cross-lingual NER Transfer Accuracy and Unlabeled Target-Language Data Volume in Low-Resource Languages

Multilingual Language Models (MLLMs) exhibit robust cross-lingual transfer capabilities, or the abil

Cross-lingual NER Performance with Unlabeled Target Data in Low-Resource Settings

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Scaling Out-of-Domain Unlabeled Data for Robust Cross-Lingual NER Alignment in Low-Resource Settings

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Cross-Lingual NER for Financial Transaction Data in Low-Resource Languages

We propose an efficient modeling framework for cross-lingual named entity recognition in semi-struct

Isomorphic Cross-lingual Embeddings for Low-Resource Languages

Cross-Lingual Word Embeddings (CLWEs) are a key component to transfer linguistic information learnt