Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Towards a Deep Understanding of Multilingual End-to-End Speech Translation

Domain:

natural language processing

Record type:

paper
Creator:
SunZhaLeiZhu
Host:avatar
In this paper, we employ Singular Value Canonical Correlation Analysis (SVCCA) to analyze representations learnt in a multilingual end-to-end speech translation model trained over 22 languages. SVCCA enables us to estimate representational similarity across languages and layers, enhancing our understanding of the functionality of multilingual speech translation and its potential connection to multilingual neural machine translation. The multilingual speech translation model is trained on the CoVoST 2 dataset in all possible directions, and we utilize LASER to extract parallel bitext data for SVCCA analysis. We derive three major findings from our analysis: (I) Linguistic similarity loses its efficacy in multilingual speech translation when the training data for a specific language is limited. (II) Enhanced encoder representations and well-aligned audio-text data significantly improve translation quality, surpassing the bilingual counterparts when the training data is not compromised. (III) The encoder representations of multilingual speech translation demonstrate superior performance in predicting phonetic features in linguistic typology prediction. With these findings, we propose that releasing the constraint of limited data for low-resource languages and subsequently combining them with linguistically related high-resource languages could offer a more effective approach for multilingual end-to-end speech translation. Accepted to Findings of EMNLP 2023

Visit

arxiv.org

Tasks

speech translationmachine translationspeech processing

Tags

Computation and Language

Similar

End-to-End Automatic Speech Translation of AudiobooksMultilingual Speech Recognition With A Single End-To-End ModelListen and Translate: A Proof of Concept for End-to-End Speech-to-Text TranslationTowards End-to-End Training of Automatic Speech Recognition for Nigerian PidginImproving End-to-End Speech Translation for the Low Resource Language Fongbe to Frenchabu14/end-to-end-speech-to-text

End-to-End Automatic Speech Translation of Audiobooks

We investigate end-to-end speech-to-text translation on a corpus of audiobooks specifically augmente

Multilingual Speech Recognition With A Single End-To-End Model

Training a conventional automatic speech recognition (ASR) system to support multiple languages is c

Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation

This paper proposes a first attempt to build an end-to-end speech-to-text translation system, which

Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin

Nigerian Pidgin remains one of the most popular languages in West Africa. With at least 75 million speakers along the West African coast, the language has spread to diasporic communities through Nigerian immigrants in England, Canada, and America, amongst others. I

Improving End-to-End Speech Translation for the Low Resource Language Fongbe to French

This study addresses the challenges of end-to-end (E2E) Speech-to-Text Translation (STT) for the low

abu14/end-to-end-speech-to-text

end to end amharic text to speech voice recognition project # Amharic Speech Recognition In this p