Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation

Domain:

natural language processing

Record type:

paperdataset
Creator:
AkkBanSamRav
Host:avatar
Despite Telugu being spoken by over 80 million people, speech translation research for this morphologically rich language remains severely underexplored. We address this gap by developing a high-quality Telugu--English speech translation benchmark from 46 hours of manually verified CSTD corpus data (30h/8h/8h train/dev/test split). Our systematic comparison of cascaded versus end-to-end architectures shows that while IndicWhisper + IndicMT achieves the highest performance due to extensive Telugu-specific training data, finetuned SeamlessM4T models demonstrate remarkable competitiveness despite using significantly less Telugu-specific training data. This finding suggests that with careful hyperparameter tuning and sufficient parallel data (potentially less than 100 hours), end-to-end systems can achieve performance comparable to cascaded approaches in low-resource settings. Our metric reliability study evaluating BLEU, METEOR, ChrF++, ROUGE-L, TER, and BERTScore against human judgments reveals that traditional metrics provide better quality discrimination than BERTScore for Telugu--English translation. The work delivers three key contributions: a reproducible Telugu--English benchmark, empirical evidence of competitive end-to-end performance potential in low-resource scenarios, and practical guidance for automatic evaluation in morphologically complex language pairs. Submitted to AACL IJCNLP 2025

Visit

arxiv.org

Tasks

machine translationspeech processingspeech translation

Tags

Computation and LanguageAudio and Speech Processing

Similar

Burushaski-English Speech Translation CorpusIgbo-English Machine Translation: An Evaluation BenchmarkAssamLegalTrans: A Parallel Corpus, Benchmark and Analysis for English-Assamese Machine Translation of Legal JudgmentsA Corpus for Amharic-English Speech Translation: The Case of Tourism DomainBENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation CorpusExtending the Fongbe to French Speech Translation Corpus: resources, models and benchmark

Burushaski-English Speech Translation Corpus

The Burushaski Speech–English Parallel Corpus is a community-driven language resource developed to s

Igbo-English Machine Translation: An Evaluation Benchmark

Although researchers and practitioners are pushing the boundaries and enhancing the capacities of NLP tools and methods, works on African languages are lagging. A lot of focus on well resourced languages such as English, Japanese, German, French, Russian, Mandarin

AssamLegalTrans: A Parallel Corpus, Benchmark and Analysis for English-Assamese Machine Translation of Legal Judgments

A Corpus for Amharic-English Speech Translation: The Case of Tourism Domain

Speech translation research for the major languages like English, Japanese and Spanish has been conducted since the 1980’s. But no attempt were made in speech translation to/from the under-resourced language like Amharic. These activities suffered from the lack of

BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus

There is a major shortage of Speech-to-Speech Translation (S2ST) datasets for high resource-to-low r

Extending the Fongbe to French Speech Translation Corpus: resources, models and benchmark