Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Sociolinguistic Bias and Language Inequality in Large Language Models

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ram
Éditeur:
Zenodo
Hôte:avatar
Large language models (LLMs) are increasingly deployed across multilingual applications, yet persistent disparities in performance and fairness raise concerns about equitable access to generative AI. This paper investigates sociolinguistic bias and language inequality in LLMs across six languages: English, French, and Mandarin Chinese as high-resource languages, and Yorùbá, Swahili, and Haitian Creole as lower-resource languages. It also examines three English-based code-switching settings: English-Yorùbá, English-Swahili, and English-Haitian Creole. Using a controlled evaluation framework comprising stereotype-eliciting prompts, neutral controls, factual question answering, creative generation, and code-switching prompts, the study measures stereotype replication, harmful generalization, fluency, factual accuracy, and reliability across repeated samples. The core empirical dataset contains 17,100 non-mitigated model completions generated from five instruction-tuned systems. Mitigation analyses were conducted as a paired evaluation on the same prompt framework and are reported separately. The analysis combines human annotation, multilingual sentence embedding association tests, and statistical hypothesis testing using logistic regression and hierarchical linear models. Results show that LLMs replicate stereotypes across all evaluated languages, with higher bias incidence and weaker mitigation effects in lower-resource languages. Performance degradation is both quantitative, through lower mean fluency and accuracy, and qualitative, through mistranslations, hallucinations, language collapse, and greater output instability. Code-switching substantially increases error rates and output variance. Prompt engineering reduces measured bias but does not close cross-lingual gaps, suggesting that disparities are structurally tied to training data, alignment coverage, and evaluation regimes rather than user-level prompting alone.

Visit

doi.org

Languages

SwahiliYoruba

Tags

large language modelssociolinguistic biaslanguage inequalitycode-switchingmultilingual NLPAI fairnesslinguistic justicenatural language processing

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright (C) 2026 Sanay Ramchandani.http://rightsstatements.org/vocab/InC/1.0/

Similaires

Investigating Bias in Bulgarian in the Context of Large Language ModelsEvaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and SwahiliAfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language ModelsAraRegBias: Evaluating Dialectal and Stereotypical Bias in Arabic Large Language Models via Multi-Component MetricsTednn493/Cross-Lingual-Bias-in-Large-Language-Models-A-Comparative-Analysis-of-English-and-Swahili

Investigating Bias in Bulgarian in the Context of Large Language Models

Abstract This paper investigates the detection and annotation of bias in Bulgarian

Evaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"

This paper evaluates the understanding and biases of large language models (LLMs) regarding

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models

Existing AI bias evaluation benchmarks largely reflect Western perspectives, leaving African context

AraRegBias: Evaluating Dialectal and Stereotypical Bias in Arabic Large Language Models via Multi-Component Metrics

Abstract Large Language Models (LLMs) have achieved strong multilingual capabiliti

Tednn493/Cross-Lingual-Bias-in-Large-Language-Models-A-Comparative-Analysis-of-English-and-Swahili

# Cross Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili ## Des