This dataset is designed for text classification in Moroccan Darija, a dialect spoken in Morocco. It has been created synthetically using the Gemini-2.0-Flash model and is intended for fine-tuning the ModernBERT model. The dataset is a translation of sentence-transformers/all-nli in Moroccan Darija.
Dataset Creation