Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Multilingual TEDx Corpus for Speech Recognition and Translation

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
SalWieBreCat
Hôte:avatar
We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source languages. We segment transcripts into sentences and align them to the source-language audio and target-language translations. The corpus is released along with open-sourced code enabling extension to new talks and languages as they become available. Our corpus creation methodology can be applied to more languages than previous work, and creates multi-way parallel evaluation sets. We provide baselines in multiple ASR and ST settings, including multilingual models to improve translation performance for low-resource language pairs. Accepted to Interspeech 2021

Visit

arxiv.org

Tasks

automatic speech recognitionmachine translationspeech processingspeech translation

Tags

Computation and Language

Similaires

Noise-Robust Multilingual Speech Recognition and the Tatar Speech CorpusCross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other LanguagesDARIJA-C: towards a Moroccan DARIJA Speech recognition and speech-to-text Translation CorpusThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionKARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

Noise-Robust Multilingual Speech Recognition and the Tatar Speech Corpus

After focusing on individual languages for a long time, multilingual automatic speech recognition ha

Cross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other Languages

In a conventional Speech emotion recognition (SER) task, a classifier for a given language is traine

DARIJA-C: towards a Moroccan DARIJA Speech recognition and speech-to-text Translation Corpus

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda. The corp

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the Un

KARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

While large-scale pre-trained models have significantly advanced multilingual Automatic Speech Recog