Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases

Domain:

natural language processing

Record type:

paper
Creator:
GuoCal
Publisher:
arXiv
Host:avatar
With the starting point that implicit human biases are reflected in the statistical regularities of language, it is possible to measure biases in English static word embeddings. State-of-the-art neural language models generate dynamic word embeddings dependent on the context in which the word appears. Current methods measure pre-defined social and intersectional biases that appear in particular contexts defined by sentence templates. Dispensing with templates, we introduce the Contextualized Embedding Association Test (CEAT), that can summarize the magnitude of overall bias in neural language models by incorporating a random-effects model. Experiments on social and intersectional biases show that CEAT finds evidence of all tested biases and provides comprehensive information on the variance of effect magnitudes of the same bias in different contexts. All the models trained on English corpora that we study contain biased representations. Furthermore, we develop two methods, Intersectional Bias Detection (IBD) and Emergent Intersectional Bias Detection (EIBD), to automatically identify the intersectional biases and emergent intersectional biases from static word embeddings in addition to measuring them in contextualized word embeddings. We present the first algorithmic bias detection findings on how intersectional group members are strongly associated with unique emergent biases that do not overlap with the biases of their constituent minority identities. IBD and EIBD achieve high accuracy when detecting the intersectional and emergent biases of African American females and Mexican American females. Our results indicate that biases at the intersection of race and gender associated with members of multiple minority groups, such as African American females and Mexican American females, have the highest magnitude across all neural language models. 19 pages, 2 figures, 4 tables

Visit

doi.orgarxiv.org

Tasks

embeddings

Tags

Computers and Society (cs.CY)Artificial Intelligence (cs.AI)Computation and Language (cs.CL)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

Creative Commons Attribution Non Commercial Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-sa/4.0/legalcode

Similar

WordBias: An Interactive Visual Tool for Discovering Intersectional Biases Encoded in Word EmbeddingsPre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emotion RecognitionAraWEAT: Multidimensional Analysis of Biases in Arabic Word EmbeddingsDo self-supervised speech models develop human-like perception biases?A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesAAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark

WordBias: An Interactive Visual Tool for Discovering Intersectional Biases Encoded in Word Embeddings

Intersectional bias is a bias caused by an overlap of multiple social factors like gender, sexuality

Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emotion Recognition

Previous work has established that a person's demographics and speech style affect how well speech p

AraWEAT: Multidimensional Analysis of Biases in Arabic Word Embeddings

Recent work has shown that distributional word vector spaces often encode human biases like sexism o

Do self-supervised speech models develop human-like perception biases?

Self-supervised models for speech processing form representational spaces without using any external

A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages

We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) for five mid-resource languages. We then compare the performance of OSCAR-based and Wik

AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark

Detecting biases in natural language understanding (NLU) for African American Vernacular English (AA