Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MAGHREB-HOF: A Large-Scale Annotated Corpus for Hate and Offensive Language Detection in Maghrebi Arabic

Domain:

natural language processing

Record type:

dataset
Creator:
BESMES
Publisher:
Zenodo
Host:avatar

MAGHREB-HOF: A Large-Scale Corpus for Hate and Offensive Language Detection in Maghrebi Arabic

MAGHREB-HOF is a large-scale annotated corpus for hate and offensive language (HOF) detection in Maghrebi Arabic social-media text. The dataset includes dialectal Arabic, Arabizi, and code-mixed Arabic–French/English content commonly observed in North African online discussions. It is designed to support reproducible benchmarking for hate speech detection in noisy dialectal environments characterized by spelling variation, dialect mixing, and transliteration.

What this record contains

This Zenodo record provides:

  • The annotated corpus of Maghrebi social-media comments (fully anonymized).

  • The annotation guidelines used during the labeling process (PDF).

  • A high-confidence class-balanced subset used for the main binary classification experiments reported in the companion paper.

  • Annotation label definitions and metadata to enable reproducible experiments.

  • The official split-index file used in the experiments (xlsx) .
  • A README file describing the official partitions (md).

The partition file supports reproducibility for both experimental tracks reported in the companion paper:

- Sparse linear models: stratified 70/30 holdout split and five-fold cross-validation where applicable.
- Neural sequence models: stratified 70/15/15 train/validation/test split.

Source and Data Collection

Comments were collected from public social-media pages and posts relevant to the Maghrebi region. The dataset focuses on topics that frequently generate online discussions and potentially hostile interactions (e.g., politics, identity issues, sports rivalries).

All released data are fully anonymized and do not include direct personal identifiers.

Annotation and Labels

Each comment is annotated for hate and offensive language.

Binary classification labels:

  • 0 — Clean / non-offensive

  • 1 — Hate or offensive content

 Fine-grained categories:

  • Clean

  • Hate

  • Insult

  • Profanity

Additional metadata columns may be included to support analysis and reproducibility. See LABELS.md and the annotation guidelines for detailed definitions.

Core Columns

Typical dataset fields include:

  • ID — unique internal comment identifier

  • Comments_origine —  raw text

  • label — binary label (0/1)

  • label_multi — multi-class label (Clean / Hate / Insult / Profanity)

  • Comments_Date

  • Annotation_Date

  • Page

  • Post

  • Hub

  • Post_Date

  • City — Obtained with pretrained model in huggin face (Ammar-alhaj-ali/arabic-MARB…)

  • Score— Obtained with pretrained model in huggin face (Ammar-alhaj-ali/arabic-MARB…)

The dialect-identification metadata were produced using the following Hugging Face model:

Ammar-alhaj-ali/arabic-MARB…

These dialect predictions are provided as descriptive metadata and should not be interpreted as manually validated dialect annotations.

Recommended Use

The dataset can be used to train and evaluate models for:

  • hate and offensive language detection

  • toxicity detection

  • dialect-robust NLP for Maghrebi Arabic

  • Arabizi processing

  • code-mixed Arabic–French/English social-media text

Typical evaluation protocols include cross-validation or fixed train/test splits as described in the companion paper.

Ethical Considerations

This dataset contains potentially harmful language, including insults and hate speech. It is released solely for research purposes to support the development of systems for detecting and mitigating harmful content.

All comments are anonymized, and direct identifiers have been removed. Users of the dataset should comply with applicable ethical guidelines and regulations governing research on social-media data.

License

Dataset license: Creative Commons Attribution 4.0 International (CC BY 4.0)

Users must provide appropriate credit and cite the companion paper and this Zenodo dataset record.

Access and Embargo

This Zenodo record is published with an embargo. Files will become publicly available on the date of publication of the companion article. Dataset metadata and DOI remain visible during the embargo period.

Related Publication

Companion article:
Word and Character Models for Maghrebi Hate and Offensive Language Detection: A Large-Scae Facebook Corpus

Dataset Statistics Dashboard

An interactive dashboard with descriptive statistics (label distribution, metadata summaries, and exploratory views) is available at:

magh-arabic-hof.duckdns.org

Note: The dashboard is provided for exploration and reporting and does not grant access to embargoed files.

Annotation Demo (Doccano)

To support peer review and demonstrate the annotation workflow, a Doccano instance is available for reviewers to explore the labeling interface on a small demo project.

Doccano login page:
http://magh-arabic-hof.duck…

Reviewer username:
provided privately to reviewers upon request.

Password: provided privately to reviewers upon request.

Contact

For questions or embargo-related access requests during peer review:

Ahmed Zoubir Messalti
Department of Computer Science
Ferhat Abbas University, Sétif 1, Algeria

Email: ahmedzoubir.messalti@univ-setif.dz

Keywords

Maghrebi Arabic; Arabizi; code-mixing; hate speech; offensive language; toxicity detection; NLP; social media corpus; Deep Learning; Machine Learning.

Visit

doi.org

Tasks

hate speech detectiontext classification

Languages

Arabic, Algerian SpokenArabic, Libyan SpokenArabic, Moroccan SpokenArabic, Tunisian SpokenHamer-BannaNdasa

Tags

Natural language processingMachine learningDeep learningArabic LanguageHate speechSentiment Analysis

Licenses

info:eu-repo/semantics/embargoedAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Arabic Facebook Corpus for Hate and Offensive Language Detection and Sentiment Analysis (MessBess)Hate Speech and Offensive Language Detection in BengaliAdversarial Evaluation of Large Language Models for Building Robust Offensive Language Detection in Moroccan ArabicHausaHate: An Expert Annotated Corpus for Hausa Hate Speech DetectionA multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languagesBenchmarking explainable offensive language detection in Somali with human-annotated rationales

Arabic Facebook Corpus for Hate and Offensive Language Detection and Sentiment Analysis (MessBess)

Arabic Facebook Corpus for Hate and Offensive Language Detection and Sentiment Analysis (MessBess)

Hate Speech and Offensive Language Detection in Bengali

Social media often serves as a breeding ground for various hateful and offensive content. Identifyin

Adversarial Evaluation of Large Language Models for Building Robust Offensive Language Detection in Moroccan Arabic

Offensive language detection is crucial for ensuring safe and inclusive digital environments. Identi

HausaHate: An Expert Annotated Corpus for Hausa Hate Speech Detection

We introduce the first expert annotated corpus of Facebook comments for Hausa hate speech detection.

A multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages

The proliferation of online offensive language necessitates the development of effective detection m

Benchmarking explainable offensive language detection in Somali with human-annotated rationales