An extensively cleaned and deduplicated text corpus specifically curated for the Algerian dialect (Darja), combining code-switching patterns between Arabic and French.
Noise, duplicate entries, and excessively short sentences have been systematically filtered out to optimize the corpus for training and evaluating language models on low-resource Algerian African dialects.
Total Rows: 1,641,779
Train Split: 1,477,601 rows