This dataset is designed for text classification in Moroccan Darija, a dialect spoken in Morocco. It has been created synthetically using the Gemini-2.0-Flash model and is intended for fine-tuning the ModernBERT model. The dataset covers a wide range of domains, making it suitable for various NLP applications in Moroccan Darija.
Dataset Creation