Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Phonology-Guided Speech-to-Speech Translation for African Languages

Domain:

natural language processing

Record type:

paper
Creator:
OchKab
Host:avatar
We present a prosody-guided framework for speech-to-speech translation (S2ST) that aligns and translates speech \emph{without} transcripts by leveraging cross-linguistic pause synchrony. Analyzing a 6{,}000-hour East African news corpus spanning five languages, we show that \emph{within-phylum} language pairs exhibit 30--40\% lower pause variance and over 3$\times$ higher onset/offset correlation compared to cross-phylum pairs. These findings motivate \textbf{SPaDA}, a dynamic-programming alignment algorithm that integrates silence consistency, rate synchrony, and semantic similarity. SPaDA improves alignment $F_1$ by +3--4 points and eliminates up to 38\% of spurious matches relative to greedy VAD baselines. Using SPaDA-aligned segments, we train \textbf{SegUniDiff}, a diffusion-based S2ST model guided by \emph{external gradients} from frozen semantic and speaker encoders. SegUniDiff matches an enhanced cascade in BLEU (30.3 on CVSS-C vs.\ 28.9 for UnitY), reduces speaker error rate (EER) from 12.5\% to 5.3\%, and runs at an RTF of 1.02. To support evaluation in low-resource settings, we also release a three-tier, transcript-free BLEU suite (M1--M3) that correlates strongly with human judgments. Together, our results show that prosodic cues in multilingual speech provide a reliable scaffold for scalable, non-autoregressive S2ST.

Visit

arxiv.org

Tasks

machine translationspeech processingspeech translation

Tags

Audio and Speech ProcessingArtificial IntelligenceComputation and Language

Similar

Contributing to Speech-to-Speech Translation for African Low-Resource Languages : Study of French-Mooré PairNaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian LanguagesDirect Speech to Speech Translation: A Reviewwiameadnane/multilingual-speech-to-speech-translation-systemAfriVox: An African benchmark dataset for Automatic Speech Translation and Speech RecognitionJONAHKYAGABA/multilingual-african-speech-translation

Contributing to Speech-to-Speech Translation for African Low-Resource Languages : Study of French-Mooré Pair

NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages

Speech translation for low-resource languages remains fundamentally limited by the scarcity of high-quality, diverse parallel speech data, a challenge that is especially pronounced in African linguistic contexts. To address this, we introduce NaijaS2ST, a parallel

Direct Speech to Speech Translation: A Review

Speech to speech translation (S2ST) is a transformative technology that bridges global communication

wiameadnane/multilingual-speech-to-speech-translation-system

A multilingual speech-to-speech translation system that detects the input language using custom GMM

AfriVox: An African benchmark dataset for Automatic Speech Translation and Speech Recognition

This project creates a benchmark dataset for evaluating Automatic Speech Translation and Speech reco

JONAHKYAGABA/multilingual-african-speech-translation

Afrivoice is a robust audio processing pipeline designed to transcribe speech from African languages