TunisianMMLU is an evaluation benchmark designed to assess large language models' (LLM) performance in Tunisian Derja, a variety of Arabic. It consists of 22,027 multiple-choice questions, translated from selected subsets of the Massive Multitask Language Understanding (MMLU), ArabicMMLU and DarijaMMLU benchmarks to measure model performance on 44 subjects in Derja.
Task Category: Multiple-choice question answering