Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
Li,CuiWanGe,
Hôte:avatar
Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation. However, existing approaches often overlook semantic commonalities across source languages, leading to biased translation performance. In this work, we propose POTSA (Parallel Optimal Transport for Speech Alignment), a new framework based on cross-lingual parallel speech pairs and Optimal Transport, designed to bridge high- and low-resource translation gaps. First, we introduce a Bias Compensation module to coarsely align initial speech representations. Second, we impose token-level OT constraints on a Q-Former using parallel pairs to establish fine-grained representation consistency. Then, we apply a layer scheduling strategy to focus OT constraints on semantically beneficial layers. Experiments on FLEURS show our method achieves SOTA performance, with +1.29 BLEU over five common languages and +2.93 BLEU on zero-shot languages, using only 10 hours of parallel speech per language.

Visit

arxiv.org

Tasks

speech translationspeech processingmachine translation

Tags

Computation and LanguageSound

Similaires

Effectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognitionGenerating Arabic text in multilingual speech-to-speech machine translation frameworkSyarotto/speech-to-text-translationabdouaziz/Semantic-Aware-Cross-Lingual-Speech-Translation-Representation-for-WolofAGRICULTURAL E-EXTENSION SERVICES: A HYBRID OF MULTILINGUAL TRANSLATION TEXT-TO-SPEECH - A FRAMEWORKCross-lingual Matryoshka Representation Learning across Speech and Text

Effectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognition

In the recent years end to end (E2E) automatic speech recognition (ASR) systems have achieved promis

Generating Arabic text in multilingual speech-to-speech machine translation framework

Syarotto/speech-to-text-translation

Code for the final project "Speech-to-Text Translation in Swahili" of LING 575C: Speech Technology f

abdouaziz/Semantic-Aware-Cross-Lingual-Speech-Translation-Representation-for-Wolof

# Semantic-Aware Cross-Lingual Speech Representation for Wolof PhD research project that aligns **W

AGRICULTURAL E-EXTENSION SERVICES: A HYBRID OF MULTILINGUAL TRANSLATION TEXT-TO-SPEECH - A FRAMEWORK

Cross-lingual Matryoshka Representation Learning across Speech and Text

Speakers of under-represented languages face both a language barrier, as most online knowledge is in