Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
LinMaoLa Hof
Hôte:avatar
Language is not monolithic. While benchmarks, including those designed for multiple languages, are often used as proxies to evaluate the performance of Large Language Models (LLMs), they tend to overlook the nuances of within-language variation and thus fail to model the experience of speakers of non-standard dialects. Focusing on African American Vernacular English (AAVE), we present the first study aimed at objectively assessing the fairness and robustness of LLMs in handling dialects across canonical reasoning tasks, including algorithm, math, logic, and integrated reasoning. We introduce ReDial (Reasoning with Dialect Queries), a benchmark containing 1.2K+ parallel query pairs in Standardized English and AAVE. We hire AAVE speakers, including experts with computer science backgrounds, to rewrite seven popular benchmarks, such as HumanEval and GSM8K. With ReDial, we evaluate widely used LLMs, including GPT, Claude, Llama, Mistral, and the Phi model families. Our findings reveal that almost all of these widely used models show significant brittleness and unfairness to queries in AAVE. Our work establishes a systematic and objective framework for analyzing LLM bias in dialectal queries. Moreover, it highlights how mainstream LLMs provide unfair service to dialect speakers in reasoning tasks, laying a critical foundation for future research. ACL 2025 main

Visit

arxiv.org

Tags

Computation and LanguageMachine Learning

Similaires

MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and TasksFairness in Multilingual Large Language Models: Addressing the Language Disparity Gap in AI SystemsSoT: Structured-of-Thought Prompting Guides Multilingual Reasoning in Large Language ModelsEvaluation of Arabic Large Language Models on Moroccan DialectMultilingual Prompt Engineering in Large Language Models: A Survey Across NLP TasksIndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks

There has been a surge in LLM evaluation research to understand LLM capabilities and limitations. Ho

Fairness in Multilingual Large Language Models: Addressing the Language Disparity Gap in AI Systems

Current Large Language Models (LLMs) exhibit significant performance disparities across languages, w

SoT: Structured-of-Thought Prompting Guides Multilingual Reasoning in Large Language Models

Recent developments have enabled Large Language Models (LLMs) to engage in complex reasoning tasks t

Evaluation of Arabic Large Language Models on Moroccan Dialect

Large Language Models (LLMs) have shown outstanding performance in many Natural Language Processing

Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks

Large language models (LLMs) have demonstrated impressive performance across a wide range of Natural

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is sti