TuniSpeech-21h is a 21-hour speech corpus specifically designed for Tunisian Arabic (Derja). It was developed to address the underrepresentation of this dialect in the landscape of Automatic Speech Recognition (ASR). The dataset is compiled from social media (YouTube and Facebook) and broadcast materials, capturing a wide range of spontaneous speech and diverse linguistic characteristics.
Feature