Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect

Domain:

natural language processing

Record type:

paperdataset
Creator:
NaoLorLou
Host:avatar
Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we propose the LinTO audio and textual datasets -- comprehensive resources that capture phonological and lexical features of Tunisian Arabic Dialect. These datasets include a variety of texts from numerous sources and real-world audio samples featuring diverse speakers and code-switching between Tunisian Arabic Dialect and English or French. By providing high-quality audio paired with precise transcriptions, the LinTO audio and textual datasets aim to provide qualitative material to build and benchmark ASR systems for the Tunisian Arabic Dialect. Keywords -- Tunisian Arabic Dialect, Speech-to-Text, Low-Resource Languages, Audio Data Augmentation

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

Arabic, Tunisian Spoken

Tags

Computation and LanguageSoundAudio and Speech Processing

Similar

assermosa/Tunisian-Arabic-Automatic-Speech-Recognition-ASR-Automatic Code-switched Academic Tunisian Arabic Speech RecognitionA New Tunisian Arabic Corpus and Benchmark for Automatic Speech RecognitionThe Automatic Recognition and Translation of Tunisian Dialect Named Entities into Modern Standard ArabicModeling Gender and Dialect Bias in Automatic Speech RecognitionLeveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition

assermosa/Tunisian-Arabic-Automatic-Speech-Recognition-ASR-

project combines multiple Tunisian speech datasets, applies audio augmentation techniques, and achie

Automatic Code-switched Academic Tunisian Arabic Speech Recognition

A New Tunisian Arabic Corpus and Benchmark for Automatic Speech Recognition

The Automatic Recognition and Translation of Tunisian Dialect Named Entities into Modern Standard Arabic

Modeling Gender and Dialect Bias in Automatic Speech Recognition

Dialect and gender-based biases have become an area of concern in language-dependent AI systems incl

Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition

Crafting an effective Automatic Speech Recognition (ASR) solution for dialects demands innovative ap