Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Fusha–Darija Evaluation Dataset: Human-Validated Phase-2 Expansion

Domaine:

natural language processing

Type de record:

dataset
Créateur:
ELB
Éditeur:
Zenodo
Hôte:avatar
Human-validated Phase-2 expansion of the Fusha–Darija evaluation project. The release contains publishable experiment materials and anonymized human annotations for E01, E02, E03, E04, E05, E08, and E09, covering lexical semantics and etymological evidence, Modern Standard Arabic–Moroccan Darija semantic pairs and comprehension tasks, cross-variety Arabic counterparts, framing, register continua, script and code-switching variation, and digital-service requests. The public reference layer preserves primary-gold, multi-reference, ambiguous, stress-condition, and excluded-from-primary statuses rather than collapsing uncertainty. The release contains 28 validated data/annotation/reference tables with 11,284 rows. Raw reviewer workbooks, reviewer identities, private communications, model outputs, and publication results are excluded. This dataset supplements and does not replace the Harvard Dataverse V1 package: doi.org. The archive is mixed-license: project-original materials and derived anonymized annotations are CC BY 4.0; selected third-party components retain CC BY-NC 4.0 or CC BY-SA 4.0 as documented at component and row level within the archive.

Visit

doi.org

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Tags

Computational SociolinguisticsLarge Language ModelsMoroccan DarijaArabic MSADialectal variationLinguistic biasStandard ArabicLanguage models evaluation

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode© 2026 Abdennacer ELBASRI.http://rightsstatements.org/vocab/InC/1.0/