Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Rejected Dialects: Biases Against African American Language in Reward Models

Domaine:

natural language processing

Type de record:

paper
Créateur:
MirAysCheDea
Hôte:avatar
Preference alignment via reward models helps build safe, helpful, and reliable large language models (LLMs). However, subjectivity in preference judgments and the lack of representative sampling in preference data collection can introduce new biases, hindering reward models' fairness and equity. In this work, we introduce a framework for evaluating dialect biases in reward models and conduct a case study on biases against African American Language (AAL) through several experiments comparing reward model preferences and behavior on paired White Mainstream English (WME) and both machine-translated and human-written AAL corpora. We show that reward models are less aligned with human preferences when processing AAL texts vs. WME ones (-4\% accuracy on average), frequently disprefer AAL-aligned texts vs. WME-aligned ones, and steer conversations toward WME, even when prompted with AAL texts. Our findings provide a targeted analysis of anti-AAL biases at a relatively understudied stage in LLM development, highlighting representational harms and ethical questions about the desired behavior of LLMs concerning AAL. Accepted to NAACL Findings 2025

Visit

arxiv.org

Tags

Computation and LanguageArtificial IntelligenceComputers and SocietyI.2.7; K.4.2

Similaires

Evaluating the Usage of African-American Vernacular English in Large Language ModelsHow Well Do Large Language Models Understand African American Language? Causes and ImplicationsExploring Bengali Religious Dialect Biases in Large Language Models with Evaluation PerspectivesAfrican American LanguageQuantifying the Bias of Transformer-Based Language Models for African American English in Masked Language ModelingCan Small-Scale Data Poisoning Exacerbate Dialect-Linked Biases in Large Language Models?

Evaluating the Usage of African-American Vernacular English in Large Language Models

In AI, most evaluations of natural language understanding tasks are conducted in standardized dialec

How Well Do Large Language Models Understand African American Language? Causes and Implications

We focus on studying large language models (LLMs) and their ability to successfully interpret Africa

Exploring Bengali Religious Dialect Biases in Large Language Models with Evaluation Perspectives

While Large Language Models (LLM) have created a massive technological impact in the past decade, al

African American Language

From birth to early adulthood, all aspects of a child's life undergo enormous development and change

Quantifying the Bias of Transformer-Based Language Models for African American English in Masked Language Modeling

International audience In the last three years we witnessed the proliferation of inno

Can Small-Scale Data Poisoning Exacerbate Dialect-Linked Biases in Large Language Models?

Style-conditioned data poisoning is identified as a covert vector for amplifying sociolinguistic bia