Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TuniSpeech-AI/TuniSpeech-21h

Domain:

natural language processing

Record type:

dataset
Creator:
Tun
Host:
TuniSpeech-21h is a 21-hour speech corpus specifically designed for Tunisian Arabic (Derja). It was developed to address the underrepresentation of this dialect in the landscape of Automatic Speech Recognition (ASR). The dataset is compiled from social media (YouTube and Facebook) and broadcast materials, capturing a wide range of spontaneous speech and diverse linguistic characteristics. Feature

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Arabic, Tunisian Spoken

Tags

tunisian-arabicderjaspeech-recognitiontunispeechasr

Licenses

cc-by-nc-sa-4.0

Similar

TuniSpeech-AI/whisper-tunisian-dialect

TuniSpeech-AI/whisper-tunisian-dialect