Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

An Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource Languages

Domain:

natural language processing

Record type:

papermodel
Creator:
Anastasopoulos, AntoniosChiDuo
Host:avatar
For many low-resource languages, spoken language resources are more likely to be annotated with translations than with transcriptions. Translated speech data is potentially valuable for documenting endangered languages or for training speech translation systems. A first step towards making use of such data would be to automatically align spoken words with their translations. We present a model that combines Dyer et al.'s reparameterization of IBM Model 2 (fast-align) and k-means clustering using Dynamic Time Warping as a distance metric. The two components are trained jointly using expectation-maximization. In an extremely low-resource scenario, our model performs significantly better than both a neural model and a strong baseline. accepted at EMNLP 2016

Visit

arxiv.org

Tasks

machine translationspeech processingspeech translation

Tags

Computation and Language

Similar

Unsupervised Language Model Adaptation for Low-Resource LanguagesMismatching-Aware Unsupervised Translation Quality Estimation For Low-Resource LanguagesA Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource LanguagesContributing to Speech-to-Speech Translation for African Low-Resource Languages : Study of French-Mooré PairSelectNoise: Unsupervised Noise Injection to Enable Zero-Shot Machine Translation for Extremely Low-resource LanguagesNaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages

Unsupervised Language Model Adaptation for Low-Resource Languages

This paper introduces a two-way neural machine translation system from Bengali to English and vice v

Mismatching-Aware Unsupervised Translation Quality Estimation For Low-Resource Languages

Translation Quality Estimation (QE) is the task of predicting the quality of machine translation (MT

A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages

Unsupervised part of speech (POS) tagging is often framed as a clustering problem, but practical tag

Contributing to Speech-to-Speech Translation for African Low-Resource Languages : Study of French-Mooré Pair

SelectNoise: Unsupervised Noise Injection to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages

In this work, we focus on the task of machine translation (MT) from extremely low-resource language

NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages

Speech translation for low-resource languages remains fundamentally limited by the scarcity of high-