DarijaVoice-Dysarthria is the first dysarthric speech corpus for Moroccan Arabic (Darija). It contains 1,071 scripted utterances (~2 hours and 5 minutes of audio) produced by a single native Darija speaker (male, age 39) with medium-severity dysarthria.
Transcripts were written prior to recording to ensure phonemic diversity and domain coverage, then manually verified by a native Darija speaker post-recording. All transcripts are provided in Arabic script.
Contents:
1,071 .wav audio files (16 kHz, mono)
A transcript file mapping each audio file to its reference text
Intended use: Automatic speech recognition (ASR) research, pathological speech processing, low-resource speech technology, and assistive technology development for Arabic dialect speakers.
Limitations: Single speaker, scripted speech only, no phoneme-level alignment, informal severity rating.