Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Dataset Package for Fusha–Darija Evaluation Tests (1–4): Materials, Model Responses, and Metrics

Domain:

natural language processing

Record type:

dataset
Creator:
ELB
Editor:
AbdELB
Publisher:
Har
Host:avatar
يوفر هذا الايداع حزمة بيانات مرافقة للدراسة، وتشمل مواد الاختبار الكاملة للاختبارات من 1 الى 4، واستجابات النماذج كاملة، ومؤشرات القياس المرتبطة بها. صممت مواد الاختبار لقياس الفروق بين العربية الفصحى والدارجة المغربية عبر ابعاد متعددة، وتشمل ازواجا نصية، وامثالا وتعابير، ومعجما مصغرا، ومجموعات تحويل لهجي، واسئلة موجهة لقياس الانحياز عبر مسارات تجريبية متوازية. كما تتضمن الحزمة ملفات توثيق منظمة تشرح بنية البيانات وحقولها، وملف قائمة محتويات، وملفات بصمات تحقق تتيح التاكد من سلامة البيانات بعد التحميل. This deposit provides a dataset package accompanying the study and includes the complete testing materials for Tests 1–4, full model responses, and associated evaluation metrics. The testing materials were designed to assess differences between Standard Arabic and Moroccan Darija across multiple dimensions, and include paired textual samples, proverbs and expressions, a small lexicon, dialectal transformation sets, and bias-oriented prompts evaluated through parallel experimental paths. The package also includes structured documentation files describing the dataset organization and fields, a manifest of all files, and checksum files to enable verification of data integrity after download. This dataset is intended to support verification and reuse of the experimental results reported in the accompanying study.

Visit

doi.orgdataverse.harvard.edu

Tasks

text classification

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Tags

Arts and HumanitiesComputer and Information ScienceArabic languageMoroccan DarijaDialectal variationLanguage models evaluationTokenization costLinguistic biasStandard ArabicDialect vs Standard language

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Fusha–Darija Evaluation Dataset: Human-Validated Phase-2 ExpansionModel evaluation metrics without optimization.Model evaluation metrics with optimization.Model performance metrics on imbalanced dataset.BounharAbdelaziz/Moroccan-Darija-Youtube-Commons-MetricsHypotheses tests for model selection.

Fusha–Darija Evaluation Dataset: Human-Validated Phase-2 Expansion

Human-validated Phase-2 expansion of the Fusha–Darija evaluation project. The release contains publi

Model evaluation metrics without optimization.

The pandemic has significantly affected many countries including the USA, UK, Asia, the Midd

Model evaluation metrics with optimization.

The pandemic has significantly affected many countries including the USA, UK, Asia, the Midd

Model performance metrics on imbalanced dataset.

Universal Health Coverage (UHC) is a global objective aimed at providing equitable access to

BounharAbdelaziz/Moroccan-Darija-Youtube-Commons-Metrics

This dataset contains evaluation metrics for various Automatic Speech Recognition (ASR) models on Mo

Hypotheses tests for model selection.

The impact of land tenure security on smallholder crop productivity in Sub-Sahara Africa (SS