Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection

Domain:

natural language processing

Record type:

paper
Creator:
SapSwaViaZho
Publisher:
arXiv
Host:avatar
The perceived toxicity of language can vary based on someone's identity and beliefs, but this variation is often ignored when collecting toxic language datasets, resulting in dataset and model biases. We seek to understand the who, why, and what behind biases in toxicity annotations. In two online studies with demographically and politically diverse participants, we investigate the effect of annotator identities (who) and beliefs (why), drawing from social psychology research about hate speech, free speech, racist beliefs, political leaning, and more. We disentangle what is annotated as toxic by considering posts with three characteristics: anti-Black language, African American English (AAE) dialect, and vulgarity. Our results show strong associations between annotator identity and beliefs and their ratings of toxicity. Notably, more conservative annotators and those who scored highly on our scale for racist beliefs were less likely to rate anti-Black language as toxic, but more likely to rate AAE as toxic. We additionally present a case study illustrating how a popular toxicity detection system's ratings inherently reflect only specific beliefs and perspectives. Our findings call for contextualizing toxicity labels in social variables, which raises immense implications for toxic language annotation and detection. NAACL 2022 Camera Ready

Visit

doi.orgarxiv.org

Tasks

hate speech detectiontext classification

Tags

Computation and Language (cs.CL)Human-Computer Interaction (cs.HC)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Challenges in Automated Debiasing for Toxic Language DetectionMitigating Racial Biases in Toxic Language Detection with an Equity-Based Ensemble FrameworkAttitudes and beliefs of parents of children with disabilities in UgandaRacial Bias in Hate Speech and Abusive Language Detection DatasetsCurrent beliefs and attitudes regarding epilepsy in MaliHAUSA TOXIC POLITICAL LANGUAGE

Challenges in Automated Debiasing for Toxic Language Detection

Biased associations have been a challenge in the development of classifiers for detecting toxic language, hindering both fairness and accuracy. As potential solutions, we investigate recently introduced debiasing methods for text classification datasets and models,

Mitigating Racial Biases in Toxic Language Detection with an Equity-Based Ensemble Framework

Recent research has demonstrated how racial biases against users who write African American English

Attitudes and beliefs of parents of children with disabilities in Uganda

Background. Little is known about the experience of carers of children with disabilities in Uganda,

Racial Bias in Hate Speech and Abusive Language Detection Datasets

Technologies for abusive language detection are being developed and applied with little consideratio

Current beliefs and attitudes regarding epilepsy in Mali

HAUSA TOXIC POLITICAL LANGUAGE