Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Direct Speech to Speech Translation: A Review

Domaine:

natural language processing

Type de record:

paper
Créateur:
SarShaJavJam
Hôte:avatar
Speech to speech translation (S2ST) is a transformative technology that bridges global communication gaps, enabling real time multilingual interactions in diplomacy, tourism, and international trade. Our review examines the evolution of S2ST, comparing traditional cascade models which rely on automatic speech recognition (ASR), machine translation (MT), and text to speech (TTS) components with newer end to end and direct speech translation (DST) models that bypass intermediate text representations. While cascade models offer modularity and optimized components, they suffer from error propagation, increased latency, and loss of prosody. In contrast, direct S2ST models retain speaker identity, reduce latency, and improve translation naturalness by preserving vocal characteristics and prosody. However, they remain limited by data sparsity, high computational costs, and generalization challenges for low-resource languages. The current work critically evaluates these approaches, their tradeoffs, and future directions for improving real time multilingual communication.

Visit

arxiv.org

Tasks

speech translationspeech processingmachine translation

Tags

Computation and LanguageSoundAudio and Speech Processing

Similaires

BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation CorpusSpeech to Speech Translation with Translatotron: A State of the Art Reviewwiameadnane/multilingual-speech-to-speech-translation-systemEvolution of Performance Metrics for Accurate Evaluation of Speech-to-Speech Translation Models: A Literature ReviewDirect English-to-Yoruba speech Translation model using Transformer with Augmented Attention MechanismPhonology-Guided Speech-to-Speech Translation for African Languages

BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus

There is a major shortage of Speech-to-Speech Translation (S2ST) datasets for high resource-to-low r

Speech to Speech Translation with Translatotron: A State of the Art Review

A cascade-based speech-to-speech translation has been considered a benchmark for a very long time, b

wiameadnane/multilingual-speech-to-speech-translation-system

A multilingual speech-to-speech translation system that detects the input language using custom GMM

Evolution of Performance Metrics for Accurate Evaluation of Speech-to-Speech Translation Models: A Literature Review

The translation of speech from a source to speech in a target language with generative artificial in

Direct English-to-Yoruba speech Translation model using Transformer with Augmented Attention Mechanism

Phonology-Guided Speech-to-Speech Translation for African Languages

We present a prosody-guided framework for speech-to-speech translation (S2ST) that aligns and transl