Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
BeuAudAnkBal
Hôte:avatar
Existing AI bias evaluation benchmarks largely reflect Western perspectives, leaving African contexts underrepresented and enabling harmful stereotypes in applications across various domains. To address this gap, we introduce AfriStereo, the first open-source African stereotype dataset and evaluation framework grounded in local socio-cultural contexts. Through community engaged efforts across Senegal, Kenya, and Nigeria, we collected 1,163 stereotypes spanning gender, ethnicity, religion, age, and profession. Using few-shot prompting with human-in-the-loop validation, we augmented the dataset to over 5,000 stereotype-antistereotype pairs. Entries were validated through semantic clustering and manual annotation by culturally informed reviewers. Preliminary evaluation of language models reveals that nine of eleven models exhibit statistically significant bias, with Bias Preference Ratios (BPR) ranging from 0.63 to 0.78 (p <= 0.05), indicating systematic preferences for stereotypes over antistereotypes, particularly across age, profession, and gender dimensions. Domain-specific models appeared to show weaker bias in our setup, suggesting task-specific training may mitigate some associations. Looking ahead, AfriStereo opens pathways for future research on culturally grounded bias evaluation and mitigation, offering key methodologies for the AI community on building more equitable, context-aware, and globally inclusive NLP technologies.

Visit

arxiv.org

Tasks

text classification

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similaires

AraRegBias: Evaluating Dialectal and Stereotypical Bias in Arabic Large Language Models via Multi-Component MetricsEvaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"{S}tereo{S}et: Measuring stereotypical bias in pretrained language modelsSociolinguistic Bias and Language Inequality in Large Language ModelsOASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QAAn MCP Server for a Heterogeneous African Studies Collection: Grounded Search and Recommendation for Large Language Models

AraRegBias: Evaluating Dialectal and Stereotypical Bias in Arabic Large Language Models via Multi-Component Metrics

Abstract Large Language Models (LLMs) have achieved strong multilingual capabiliti

Evaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"

This paper evaluates the understanding and biases of large language models (LLMs) regarding

{S}tereo{S}et: Measuring stereotypical bias in pretrained language models

A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or African Americans are athletic. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real-world d

Sociolinguistic Bias and Language Inequality in Large Language Models

Large language models (LLMs) are increasingly deployed across multilingual applications, yet persist

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA),

An MCP Server for a Heterogeneous African Studies Collection: Grounded Search and Recommendation for Large Language Models

This package contains the data set and code used for the quantitative evaluation of the MCP server.