Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Dialect in the machine: auditing sentiment analysis models for African American Vernacular English bias

Domain:

natural language processing

Record type:

paperdataset
Creator:
Kri
Publisher:
Sci
Host:
Sentiment analysis systems are deployed at scale in content moderation, social listening, and public health, yet their behaviour on African American Vernacular English (AAVE) is rarely evaluated in a systematic way. We present a reproducible audit framework, implemented as an open Google Colab notebook, that applies three widely-used transformer models (BERT-base, RoBERTa-base, and DistilBERT) to a curated parallel corpus of 1,200 sentence pairs. Each AAVE sentence is matched to a semantically equivalent Standard American English (SAE) sentence. Across all three models, mean predicted negative-sentiment probability is 14.3 percentage points higher for AAVE text than for matched SAE text (p < 0.001, Wilcoxon signed-rank). The disparity is largest for positive-polarity sentences (+18.7 pp), which suggests that AAVE affirmative and emphatic constructions are being read as negative by the models. We further show that this gap is not explained by lexical overlap with hate-speech training data; it is instead driven by morphosyntactic AAVE features, chief among them copula deletion, negative concord, and habitual 'be'. We propose three bias-quantification metrics, the Dialect Sentiment Gap (DSG), the Polarity Flip Rate (PFR), and the Feature-Conditioned Disparity (FCD), and release all code, data, and model outputs publicly.

Visit

doi.org

Tasks

sentiment analysistext classification

Licenses

http://creativecommons.org/licenses/by/4.0/

Similar

Dialect-Specific Models for Automatic Speech Recognition of African American Vernacular EnglishAfrican American Vernacular English as a Literary DialectEvaluating the Usage of African-American Vernacular English in Large Language ModelsAfrican American Vernacular English in CaliforniaThe Origins of African American Vernacular EnglishToward a Description of African American Vernacular English Dialect Regions Using “Black Twitter”

Dialect-Specific Models for Automatic Speech Recognition of African American Vernacular English

African American Vernacular English (AAVE) is a widely-spoken dialect of English, yet it is under-represented in major speech corpora. As a result, speakers of this dialect are often misunderstood by NLP applications. This study explores the effect on transcription

African American Vernacular English as a Literary Dialect

Knowledge about one’s linguistic background, especially when it is different from mainstream varieti

Evaluating the Usage of African-American Vernacular English in Large Language Models

In AI, most evaluations of natural language understanding tasks are conducted in standardized dialec

African American Vernacular English in California

The Origins of African American Vernacular English

Toward a Description of African American Vernacular English Dialect Regions Using “Black Twitter”

Recent research has established that African American Vernacular English (AAVE) is not monolithic. H