A randomized Moroccan Darija phrase-pair dataset with three script variants per phrase: Arabic, Arabizi (Latin), and Mixte (mixed script).
This dataset is a post-processed version of the generation output. Segment order has been randomized per source to prevent models from relying on positional/sequential patterns during training.
Column
Description
video_id
Source YouTube video ID (empty for articles)
article_id