Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
JhaPraDenLas
Hôte:avatar
Recent studies have shown that Text-to-Image (T2I) model generations can reflect social stereotypes present in the real world. However, existing approaches for evaluating stereotypes have a noticeable lack of coverage of global identity groups and their associated stereotypes. To address this gap, we introduce the ViSAGe (Visual Stereotypes Around the Globe) dataset to enable the evaluation of known nationality-based stereotypes in T2I models, across 135 nationalities. We enrich an existing textual stereotype resource by distinguishing between stereotypical associations that are more likely to have visual depictions, such as `sombrero', from those that are less visually concrete, such as 'attractive'. We demonstrate ViSAGe's utility through a multi-faceted evaluation of T2I generations. First, we show that stereotypical attributes in ViSAGe are thrice as likely to be present in generated images of corresponding identities as compared to other attributes, and that the offensiveness of these depictions is especially higher for identities from Africa, South America, and South East Asia. Second, we assess the stereotypical pull of visual depictions of identity groups, which reveals how the 'default' representations of all identity groups in ViSAGe have a pull towards stereotypical depictions, and that this pull is even more prominent for identity groups from the Global South. CONTENT WARNING: Some examples contain offensive stereotypes. Association for Computational Linguistics (ACL) 2024

Visit

arxiv.org

Tasks

computer visionimage-text retrieval

Tags

Computer Vision and Pattern RecognitionComputation and LanguageComputers and Society

Similaires

Image-to-Image Translation Approach for Page Layout Analysis and Artificial Generation of Historical ManuscriptsThe Analysis of a GPT-based Sepedi Text Generation ModelText Image Generation for Low-Resource Languages with Dual Translation LearningGoing PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global SouthAltDiffusion: A Multilingual Text-to-Image Diffusion ModelA transformer-based approach to Nigerian Pidgin text generation

Image-to-Image Translation Approach for Page Layout Analysis and Artificial Generation of Historical Manuscripts

Preprint version International audience Document layout analysis is essential in Opti

The Analysis of a GPT-based Sepedi Text Generation Model

Text generation is defined as a component of natural language processing that makes use of computa

Text Image Generation for Low-Resource Languages with Dual Translation Learning

Scene text recognition in low-resource languages frequently faces challenges due to the limited avai

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South

Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely cal

AltDiffusion: A Multilingual Text-to-Image Diffusion Model

Large Text-to-Image(T2I) diffusion models have shown a remarkable capability to produce photorealist

A transformer-based approach to Nigerian Pidgin text generation

Abstract This paper describes the development of a transformer-based text generation model for Nige