Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Multimodal In-context Learning for ASR of Low-resource Languages

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
Li,Nie
Hôte:avatar
Automatic speech recognition (ASR) still covers only a small fraction of the world's languages, mainly due to supervised data scarcity. In-context learning (ICL) with large language models (LLMs) addresses this problem, but prior work largely focuses on high-resource languages covered during training and text-only settings. This paper investigates whether speech LLMs can learn unseen languages with multimodal ICL (MICL), and how this learning can be used to improve ASR. We conduct experiments with two speech LLMs, Phi-4 and Qwen3-Omni, on three diverse endangered languages. Firstly, we find that MICL is effective for unseen languages, leveraging both speech and text modalities. We further show that cross-lingual transfer learning improves MICL efficiency on target languages without training on them. Moreover, we analyze attention patterns to interpret MICL mechanisms, and we observe layer-dependent preferences between audio and text context, with an overall bias towards text. Finally, we show that prompt-based ASR with speech LLMs performs poorly on unseen languages, motivating a simple ASR system that combines a stronger acoustic model with a speech LLM via MICL-based selection of acoustic hypotheses. Results show that MICL consistently improves ASR performance, and that cross-lingual transfer learning matches or outperforms corpus-trained language models without using target-language data. Our code is publicly available. ACL 2026 findings

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similaires

elerdg/ASR-for-low-resource-languagesMultimodal Teacher-Student Learning for Cross-Lingual Entity Recognition in Low-Resource LanguagesFeature learning for efficient ASR-free keyword spotting in low-resource languagesOptimal Transport Distillation for Multimodal Retrieval in Low-Resource LanguagesLarge Multimodal Models for Low-Resource Languages: A SurveyRafat-decodis/Robust-ASR-for-Low-Resource-Languages

elerdg/ASR-for-low-resource-languages

Fine-tune wav2vec2-xls-r on data from low-resource-languages # ASR for Low-resource languages ## O

Multimodal Teacher-Student Learning for Cross-Lingual Entity Recognition in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Feature learning for efficient ASR-free keyword spotting in low-resource languages

We consider feature learning for efficient keyword spotting that can be applied in severely under-re

Optimal Transport Distillation for Multimodal Retrieval in Low-Resource Languages

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Large Multimodal Models for Low-Resource Languages: A Survey

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) fo

Rafat-decodis/Robust-ASR-for-Low-Resource-Languages

Exploring Benchmark Gaps and Real-World Speech Generalization for Language in Low Resource # 🧠 A Ro