Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Dataset Package for Fusha–Darija Evaluation Tests (1–4): Materials, Model Responses, and Metrics

Domaine:

natural language processing

Type de record:

dataset
Créateur:
ELB
Éditeur:
AbdELB
Éditeur:
Har
Hôte:avatar
يوفر هذا الايداع حزمة بيانات مرافقة للدراسة، وتشمل مواد الاختبار الكاملة للاختبارات من 1 الى 4، واستجابات النماذج كاملة، ومؤشرات القياس المرتبطة بها. صممت مواد الاختبار لقياس الفروق بين العربية الفصحى والدارجة المغربية عبر ابعاد متعددة، وتشمل ازواجا نصية، وامثالا وتعابير، ومعجما مصغرا، ومجموعات تحويل لهجي، واسئلة موجهة لقياس الانحياز عبر مسارات تجريبية متوازية. كما تتضمن الحزمة ملفات توثيق منظمة تشرح بنية البيانات وحقولها، وملف قائمة محتويات، وملفات بصمات تحقق تتيح التاكد من سلامة البيانات بعد التحميل. This deposit provides a dataset package accompanying the study and includes the complete testing materials for Tests 1–4, full model responses, and associated evaluation metrics. The testing materials were designed to assess differences between Standard Arabic and Moroccan Darija across multiple dimensions, and include paired textual samples, proverbs and expressions, a small lexicon, dialectal transformation sets, and bias-oriented prompts evaluated through parallel experimental paths. The package also includes structured documentation files describing the dataset organization and fields, a manifest of all files, and checksum files to enable verification of data integrity after download. This dataset is intended to support verification and reuse of the experimental results reported in the accompanying study.

Visit

doi.orgdataverse.harvard.edu

Tasks

text classification

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Tags

Arts and HumanitiesComputer and Information ScienceArabic languageMoroccan DarijaDialectal variationLanguage models evaluationTokenization costLinguistic biasStandard ArabicDialect vs Standard language

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Fusha–Darija Evaluation Dataset: Human-Validated Phase-2 ExpansionModel evaluation metrics without optimization.Model evaluation metrics with optimization.Model performance metrics on imbalanced dataset.BounharAbdelaziz/Moroccan-Darija-Youtube-Commons-MetricsHypotheses tests for model selection.

Fusha–Darija Evaluation Dataset: Human-Validated Phase-2 Expansion

Human-validated Phase-2 expansion of the Fusha–Darija evaluation project. The release contains publi

Model evaluation metrics without optimization.

The pandemic has significantly affected many countries including the USA, UK, Asia, the Midd

Model evaluation metrics with optimization.

The pandemic has significantly affected many countries including the USA, UK, Asia, the Midd

Model performance metrics on imbalanced dataset.

Universal Health Coverage (UHC) is a global objective aimed at providing equitable access to

BounharAbdelaziz/Moroccan-Darija-Youtube-Commons-Metrics

This dataset contains evaluation metrics for various Automatic Speech Recognition (ASR) models on Mo

Hypotheses tests for model selection.

The impact of land tenure security on smallholder crop productivity in Sub-Sahara Africa (SS