Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Detection of Offensive Language and ITS Severity for Low Resource Language

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
RamHamSadNai
Éditeur:
Ass
Hôte:
Continuous proliferation of hate speech in different languages on social media has drawn significant attention from researchers in the past decade. Detecting hate speech is indispensable irrespective of the scale of use of language, as it inflicts huge harm on society. This work presents a first resource for classifying the severity of hate speech in addition to classifying offensive and hate speech content. Current research mostly limits hate speech classification to its primary categories, such as racism, sexism, and hatred of religions. However, hate speech targeted at different protected characteristics also manifests in different forms and intensities. It is important to understand varying severity levels of hate speech so that the most harmful cases of hate speech may be identified and dealt with earlier than the less harmful ones. In this work, we focus on detecting offensive speech, hate speech, and multiple levels of hate speech in the Urdu language. We investigate three primary target categories of hate speech: religion, racism, and national origin. We further divide these categories into levels based on the severity of hate conveyed. The severity levels are referred to as symbolization , insult , and attribution . A corpus comprising more than 20,000 tweets against the corresponding hate speech categories and severity levels is collected and annotated. A comprehensive experimentation scheme is applied using traditional as well as deep learning–based models to examine their impact on hate speech detection. The highest macro-averaged F-score yielded for detecting offensive speech is 86% while the highest F-scores for detecting hate speech with respect to ethnicity, national origin, and religious affiliation are 80%, 81%, and 72%, respectively. This shows that results are very encouraging and would provide a lead towards further investigation in this domain.

Visit

doi.org

Tasks

hate speech detectiontext classification

Licenses

https://www.acm.org/publications/policies/copyright_policy#Background

Similaires

A Forensic Linguistic Dataset for Offensive Content Detection in Low-Resource Language: SetswanaInvestigating Offensive Language Detection in a Low-Resource Setting with a Robustness PerspectiveCross-lingual Offensive Language Identification for Low Resource Languages: The Case of MarathiOffensive Language Detection in ArabizirayenFathallah/Tunisian-Offensive-Language-DetectionHichamDe/darija-offensive-language-detection

A Forensic Linguistic Dataset for Offensive Content Detection in Low-Resource Language: Setswana

Developing Monolingual Setswana Datasets for Offensive Content Detection Reproducibility Package, Me

Investigating Offensive Language Detection in a Low-Resource Setting with a Robustness Perspective

Moroccan Darija, a dialect of Arabic, presents unique challenges for natural language processing due

Cross-lingual Offensive Language Identification for Low Resource Languages: The Case of Marathi

The widespread presence of offensive language on social media motivated the development of systems c

Offensive Language Detection in Arabizi

Detecting offensive language in under-resourced languages presents a significant real-world challeng

rayenFathallah/Tunisian-Offensive-Language-Detection

This repository contains the implementation of a real-time content filtering system designed to dete

HichamDe/darija-offensive-language-detection

# 🚀 Offensive Message Detection in Arabic Darija This project focuses on building a **classificatio