Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark

Domain:

natural language processing

Record type:

paperdataset
Creator:
GupMenYurO’b
Host:avatar
Detecting biases in natural language understanding (NLU) for African American Vernacular English (AAVE) is crucial to developing inclusive natural language processing (NLP) systems. To address dialect-induced performance discrepancies, we introduce AAVENUE ({AAVE} {N}atural Language {U}nderstanding {E}valuation), a benchmark for evaluating large language model (LLM) performance on NLU tasks in AAVE and Standard American English (SAE). AAVENUE builds upon and extends existing benchmarks like VALUE, replacing deterministic syntactic and morphological transformations with a more flexible methodology leveraging LLM-based translation with few-shot prompting, improving performance across our evaluation metrics when translating key tasks from the GLUE and SuperGLUE benchmarks. We compare AAVENUE and VALUE translations using five popular LLMs and a comprehensive set of metrics including fluency, BARTScore, quality, coherence, and understandability. Additionally, we recruit fluent AAVE speakers to validate our translations for authenticity. Our evaluations reveal that LLMs consistently perform better on SAE tasks than AAVE-translated versions, underscoring inherent biases and highlighting the need for more inclusive NLP models. We have open-sourced our source code on GitHub and created a website to showcase our work at aavenuee.github.io. Published at NLP4PI @ EMNLP 2024

Visit

arxiv.org

Tags

Computation and Language

Similar

Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like BiasesKambaata LLM Cultural Benchmarkrifaasa/hassaniya-llm-benchmarkFatika01/nigeria-livestock-llm-benchmarkBehailuBerhanu/kambaata-llm-cultural-benchmarkFatikafarouq/llm-nigeria-livestock-benchmark

Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases

With the starting point that implicit human biases are reflected in the statistical regularities of

Kambaata LLM Cultural Benchmark

Maintenance release for Zenodo archival of the Kambaata LLM Cultural Benchmark. This release contain

rifaasa/hassaniya-llm-benchmark

benchmark of modern LLMs on Hassaniya Arabic dialect » # hassaniya-llm-benchmark Code and evaluati

Fatika01/nigeria-livestock-llm-benchmark

A 420-question benchmark evaluating LLM performance on Nigerian livestock management knowledge, desi

BehailuBerhanu/kambaata-llm-cultural-benchmark

A 77-item benchmark for evaluating cultural knowledge, hallucination, and epistemic behavior in larg

Fatikafarouq/llm-nigeria-livestock-benchmark

# LLM Nigeria Livestock Benchmark Benchmarking LLM performance on Nigerian indigenous livestock man