Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

XHate-999: Analyzing and Detecting Abusive Language Across Domains and Languages

Domaine:

natural language processing

Type de record:

dataset
Créateur:
GlaKarVul
Éditeur:
ApoUni
Éditeur:
Int
Hôte:avatar
We present XHate -999, a multi-domain and multilingual evaluation data set for abusive language detection. By aligning test instances across six typologically diverse languages, XHate-999 for the first time allows for disentanglement of the domain transfer and language transfer effects in abusive language detection. We conduct a series of domain- and language-transfer experiments with state-of-the-art monolingual and multilingual transformer models, setting strong baseline results and profiling XH ATE -999 as a comprehensive evaluation resource for abusive language detection. Finally, we show that domain- and language-adaptation, via intermediate masked language modeling on abusive corpora in the target language, can lead to substantially improved abusive language detection in the target language in the zero-shot transfer setups.

Visit

doi.orgwww.repository.cam.ac.uk

Tasks

hate speech detectiontext classificationtransfer learning

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeopen.accesshttp://purl.org/coar/access_right/c_abf2

Similaires

In What Languages are Generative Language Models the Most Formal? Analyzing Formality Distribution across LanguagesDETECTING CYBERBULLYING ACROSS NIGERIAN LANGUAGES: A MULTILINGUAL SYSTEMAfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African LanguagesAdvances in Amazigh Language Technologies: A Comprehensive Survey Across Processing DomainsAnalyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal RepresentationsIntersectional Bias in Hate Speech and Abusive Language Datasets

In What Languages are Generative Language Models the Most Formal? Analyzing Formality Distribution across Languages

Multilingual generative language models (LMs) are increasingly fluent in a large variety of language

DETECTING CYBERBULLYING ACROSS NIGERIAN LANGUAGES: A MULTILINGUAL SYSTEM

ABSTRACT

As the digital world evolves, the

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

Hate speech and abusive language are global phenomena that need socio-cultural background knowledge

Advances in Amazigh Language Technologies: A Comprehensive Survey Across Processing Domains

The Amazigh language, spoken by millions across North Africa, presents unique computational challeng

Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations

While understanding the knowledge boundaries of LLMs is crucial to prevent hallucination, research o

Intersectional Bias in Hate Speech and Abusive Language Datasets

Algorithms are widely applied to detect hate speech and abusive language in social media. We investi