Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English

Domain:

natural language processing

Record type:

paperdataset
Creator:
BouMdhEllEst
Host:avatar
In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the data scarcity obstacle for a number of Arabic dialects. We collected, segmented, transcribed and translated 108 TEDx talks following our internally developed annotations guidelines. The collected talks represent 25 hours of speech with code-switching that cover speakers with various accents from over 11 different regions of Tunisia. We make the annotation guidelines and corpus publicly available. This will enable the extension of TEDxTN to new talks as they become available. We also report results for strong baseline systems of Speech Recognition and Speech Translation using multiple pre-trained and fine-tuned end-to-end models. This corpus is the first open source and publicly available speech translation corpus of Code-Switching Tunisian dialect. We believe that this is a valuable resource that can motivate and facilitate further research on the natural language processing of Tunisian Dialect. The Third Arabic Natural Language Processing Conference. Association for Computational Linguistics. 2025

Visit

arxiv.org

Tasks

automatic speech recognitioncode switchingmachine translationspeech processingspeech translation

Languages

Arabic, Tunisian Spoken

Tags

Computation and LanguageArtificial Intelligence

Similar

ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic-English ArzEn - ST: مجموعة ترجمة الكلام ثلاثية الاتجاهات لتبديل الرموز المصرية العربية- الإنجليزية ArzEn-ST : A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic-English ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic-EnglishGhana English-Twi Code-switched Speech CorpusAutomatic Code-switched Academic Tunisian Arabic Speech RecognitionNeural-based NLP systems for code-switched Arabic-English speechTunSwitch: A comprehensive Tunisian Arabic code-switched text-to-speech dataseLeveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition

ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic-English ArzEn - ST: مجموعة ترجمة الكلام ثلاثية الاتجاهات لتبديل الرموز المصرية العربية- الإنجليزية ArzEn-ST : A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic-English ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic-English

We present our work on collecting ArzEn-ST, a code-switched Egyptian Arabic -English Speech Translat

Ghana English-Twi Code-switched Speech Corpus

Gold-standard English-Twi code-switched speech with transcripts and speaker meta

Automatic Code-switched Academic Tunisian Arabic Speech Recognition

Neural-based NLP systems for code-switched Arabic-English speech

In the ever-evolving language landscape, code-switching has emerged as an interesting linguistic phe

TunSwitch: A comprehensive Tunisian Arabic code-switched text-to-speech datase

@misc{abdallah2023leveraging, title={Leveraging Data Collection and Unsupervised Learning for Code-s

Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition

Crafting an effective Automatic Speech Recognition (ASR) solution for dialects demands innovative ap