Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

The Multilingual TEDx Corpus for Speech Recognition and Translation

Domain:

natural language processing

Record type:

paperdataset
Creator:
SalWieBreCat
Host:avatar
We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source languages. We segment transcripts into sentences and align them to the source-language audio and target-language translations. The corpus is released along with open-sourced code enabling extension to new talks and languages as they become available. Our corpus creation methodology can be applied to more languages than previous work, and creates multi-way parallel evaluation sets. We provide baselines in multiple ASR and ST settings, including multilingual models to improve translation performance for low-resource language pairs. Accepted to Interspeech 2021

Visit

arxiv.org

Tasks

automatic speech recognitionmachine translationspeech processingspeech translation

Tags

Computation and Language

Similar

Noise-Robust Multilingual Speech Recognition and the Tatar Speech CorpusCross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other LanguagesDARIJA-C: towards a Moroccan DARIJA Speech recognition and speech-to-text Translation CorpusThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionKARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

Noise-Robust Multilingual Speech Recognition and the Tatar Speech Corpus

After focusing on individual languages for a long time, multilingual automatic speech recognition ha

Cross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other Languages

In a conventional Speech emotion recognition (SER) task, a classifier for a given language is traine

DARIJA-C: towards a Moroccan DARIJA Speech recognition and speech-to-text Translation Corpus

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda. The corp

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the Un

KARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

While large-scale pre-trained models have significantly advanced multilingual Automatic Speech Recog