Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
Li,LiuLiuNgu
Hôte:avatar
This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translation (MT) models, and end-to-end (E2E) Speech Translation (ST) systems for three language pairs: Bemba, North Levantine Arabic, and Tunisian Arabic into English. Building upon pre-trained models, we fine-tune our systems with different strategies to utilize resources efficiently. This study further explores system enhancement with synthetic data and model regularization. Specifically, we investigate MT-augmented ST by generating translations from ASR data using MT models. For North Levantine, which lacks parallel ST training data, a system trained solely on synthetic data slightly surpasses the cascaded system trained on real data. We also explore augmentation using text-to-speech models by generating synthetic speech from MT data, demonstrating the benefits of synthetic data in improving both ASR and ST performance for Bemba. Additionally, we apply intra-distillation to enhance model performance. Our experiments show that this approach consistently improves results across ASR, MT, and ST tasks, as well as across different pre-trained models. Finally, we apply Minimum Bayes Risk decoding to combine the cascaded and end-to-end systems, achieving an improvement of approximately 1.5 BLEU points.

Visit

arxiv.org

Tasks

automatic speech recognitionmachine translationspeech processingspeech translation

Languages

Arabic, Tunisian SpokenBemba

Tags

Computation and LanguageArtificial Intelligence

Similaires

Semi-supervised Neural Machine Translation with Consistency Regularization for Low-Resource LanguagesIMS' Systems for the IWSLT 2021 Low-Resource Speech Translation TaskSynthetic Data Diversity vs. Back-Translation for Multilingual NER in Low-Resource LanguagesGMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared TaskLIA and ELYADATA systems for the IWSLT 2025 low-resource speech translation shared taskAn Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource Languages

Semi-supervised Neural Machine Translation with Consistency Regularization for Low-Resource Languages

The advent of deep learning has led to a significant gain in machine translation. However, most of t

IMS' Systems for the IWSLT 2021 Low-Resource Speech Translation Task

This paper describes the submission to the IWSLT 2021 Low-Resource Speech Translation Shared Task by

Synthetic Data Diversity vs. Back-Translation for Multilingual NER in Low-Resource Languages

Named Entity Recognition(NER) for low-resource languages aims to produce robust systems for language

GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task

This paper describes the GMU systems for the IWSLT 2025 low-resource speech translation shared task.

LIA and ELYADATA systems for the IWSLT 2025 low-resource speech translation shared task

International audience

In this paper, we present the approach and system setu

An Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource Languages

For many low-resource languages, spoken language resources are more likely to be annotated with tran