Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
Du,PanMa,Yan
Hôte:avatar
Multimodal Large Language Models (MLLMs) have achieved significant success in Speech-to-Text Translation (S2TT) tasks. While most existing research has focused on English-centric translation directions, the exploration of many-to-many translation is still limited by the scarcity of parallel data. To address this, we propose a three-stage curriculum learning strategy that leverages the machine translation capabilities of large language models and adapts them to S2TT tasks, enabling effective learning in low-resource settings. We trained MLLMs with varying parameter sizes (3B, 7B, and 32B) and evaluated the proposed strategy using the FLEURS and CoVoST-2 datasets. Experimental results show that the proposed strategy achieves state-of-the-art average performance in $15\times14$ language pairs, requiring fewer than 10 hours of speech data per language to achieve competitive results. The source code and models are released at github.com. Accepted in ACL 2025 (Main)

Visit

arxiv.org

Tasks

machine translationspeech processingspeech translation

Tags

Computation and Language

Similaires

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech EncodersBreaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech EncodersToucan: Many-to-Many Translation for 150 African Language PairsFixing Rogue Memorization in Many-to-One Multilingual Translators of Extremely-Low-Resource Languages by Rephrasing Training SamplesLeveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense RetrievalMany-to-English Machine Translation Tools, Data, and Pretrained Models

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

Multimodal large language models (MLLMs) have achieved significant success in speech-to-text transla

Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders

Multimodal large language models (MLLMs) have achieved significant success in speech-to-text transla

Toucan: Many-to-Many Translation for 150 African Language Pairs

We address a notable gap in Natural Language Processing (NLP) by introducing a collection of resourc

Fixing Rogue Memorization in Many-to-One Multilingual Translators of Extremely-Low-Resource Languages by Rephrasing Training Samples

In this paper we study the fine-tuning of pre-trained large high-resource language models (LLMs) int

Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval

There has been limited success for dense retrieval models in multilingual retrieval, due to uneven a

Many-to-English Machine Translation Tools, Data, and Pretrained Models

While there are more than 7000 languages in the world, most translation research efforts have target