Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
LucMurAl-Uch
Hôte:avatar
Harmful content detectors-particularly disinformation classifiers-are predominantly developed and evaluated on Standard American English (SAE), leaving their robustness to dialectal variation unexplored. We present DIA-HARM, the first benchmark for evaluating disinformation detection robustness across 50 English dialects spanning U.S., British, African, Caribbean, and Asia-Pacific varieties. Using Multi-VALUE's linguistically grounded transformations, we introduce D3 (Dialectal Disinformation Detection), a corpus of 195K samples derived from established disinformation benchmarks. Our evaluation of 16 detection models reveals systematic vulnerabilities: human-written dialectal content degrades detection by 1.4-3.6% F1, while AI-generated content remains stable. Fine-tuned transformers substantially outperform zero-shot LLMs (96.6% vs. 78.3% best-case F1), with some models exhibiting catastrophic failures exceeding 33% degradation on mixed content. Cross-dialectal transfer analysis across 2,450 dialect pairs shows that multilingual models (mDeBERTa: 97.2% average F1) generalize effectively, while monolingual models like RoBERTa and XLM-RoBERTa fail on dialectal inputs. These findings demonstrate that current disinformation detectors may systematically disadvantage hundreds of millions of non-SAE speakers worldwide. We release the DIA-HARM framework, D3 corpus, and evaluation tools: github.com Accepted to ACL 2026

Visit

arxiv.org

Tasks

hate speech detectiontext classification

Tags

Computation and Language

Similaires

abrhaleyarefaine1997/meta-tigrinya-harmful-content-detectorResources Building for Arabic Harmful Online Content: SurveyUsing language sample analyses across English dialects: A case-based approach for preschoolersINTERNET REPRESENTATIONS OF DIALECTAL ENGLISHProcessing focus and accent across dialects.The Paradox of Undetected Harm: Content Moderation Blind Spots in Low-Resource Languages

abrhaleyarefaine1997/meta-tigrinya-harmful-content-detector

Machine learning-based system for detecting harmful content (hate speech, offensive language, misinf

Resources Building for Arabic Harmful Online Content: Survey

International audience

Users of social networks and Internet sites face numer

Using language sample analyses across English dialects: A case-based approach for preschoolers

This study compared language samples from typically developing 4-year-olds who spoke African Amer

INTERNET REPRESENTATIONS OF DIALECTAL ENGLISH

International audience This paper presents an account of how alternative spellings fo

Processing focus and accent across dialects.

International audience Previous research has shown that native listeners of British E

The Paradox of Undetected Harm: Content Moderation Blind Spots in Low-Resource Languages

This paper explores a systemic paradox in global content moderation: harmful content in low-resource