Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

DziriOFN Corpus (Dziri Offensive corpus) v1.0

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Oussama BoucheritKheireddine Abainia

Dziri refers to the name of the Algerian dialectal Arabic. DziriOFN is a new corpus dedicated to offensive language detection on this under-resourced language. Dziri dialect if known as a complex socio-linguistic situation, where the latter is known by the code-switching phenomenon and borrowed words from several languages (e.g. Turkish, Berber, French, Spanish, etc.). The corpus is crawled from Facebook social media, because the latter is the common used social media network by the Algerian community. In overall, the corpus contains 8.7k texts manually annotated as normal, offensive and abusive, where it was annotated by five different native speakers. The texts are written in Arabic script, and some of them are written in Latin script (i.e. Arabizi).

Table 1. DziriOFN corpus description

#texts
offensive texts3,227
abusive texts1,334
normal texts4,188
total8,749

Visit

github.com

Connected records

paper

Tasks

sentiment analysishate speech detectiontext classification

Languages

Arabic, Algerian SpokenMòoré

Tags

DziriOFNAlgerian dialectal Arabicoffensive language detection

Similaires

Zomi Monolingual Corpus v1.0Arabic Facebook Corpus for Hate and Offensive Language Detection and Sentiment Analysis (MessBess)xprogramer/DziriOFNOffensive Language Detection in Code-Mixed Bambara-French Corpus: Evaluating machine learning and deep learning classifiersMAGHREB-HOF: A Large-Scale Annotated Corpus for Hate and Offensive Language Detection in Maghrebi ArabicAAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-Corpus

Zomi Monolingual Corpus v1.0

A reviewed monolingual corpus of 363,401 Zomi (Tedim Chin, ISO 639-3 ctd) sentences with permanent i

Arabic Facebook Corpus for Hate and Offensive Language Detection and Sentiment Analysis (MessBess)

Arabic Facebook Corpus for Hate and Offensive Language Detection and Sentiment Analysis (MessBess)

xprogramer/DziriOFN

The corpus for offensive language detection in under-resourced Algerian dialectal Arabic # DziriOFN

Offensive Language Detection in Code-Mixed Bambara-French Corpus: Evaluating machine learning and deep learning classifiers

In this paper, we deal with offensive and abusive language detection on Bambara language, which is an under-resourced language mainly spoken in Mali and some other African countries. As a first work on this language, we aim to release OBAM v1.0 corpus compiling 4k

MAGHREB-HOF: A Large-Scale Annotated Corpus for Hate and Offensive Language Detection in Maghrebi Arabic

MAGHREB-HOF: A Large-Scale Corpus for Hate and Offensive Language Detection in Maghrebi Ar

AAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-Corpus

Dataset and code for AAAC: Algerian Arabic Adversarial Corpus for dialect-aware hate speech and prom