Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Decomposed evaluations of geographic disparities in text-to-image models

Domaine:

natural language processing

Type de record:

paper
Créateur:
SurPadPerSah
Hôte:avatar
Recent work has identified substantial disparities in generated images of different geographic regions, including stereotypical depictions of everyday objects like houses and cars. However, existing measures for these disparities have been limited to either human evaluations, which are time-consuming and costly, or automatic metrics evaluating full images, which are unable to attribute these disparities to specific parts of the generated images. In this work, we introduce a new set of metrics, Decomposed Indicators of Disparities in Image Generation (Decomposed-DIG), that allows us to separately measure geographic disparities in the depiction of objects and backgrounds in generated images. Using Decomposed-DIG, we audit a widely used latent diffusion model and find that generated images depict objects with better realism than backgrounds and that backgrounds in generated images tend to contain larger regional disparities than objects. We use Decomposed-DIG to pinpoint specific examples of disparities, such as stereotypical background generation in Africa, struggling to generate modern vehicles in Africa, and unrealistically placing some objects in outdoor settings. Informed by our metric, we use a new prompting structure that enables a 52% worst-region improvement and a 20% average improvement in generated background diversity.

Visit

arxiv.org

Tags

Computer Vision and Pattern RecognitionArtificial IntelligenceComputers and SocietyMachine Learning

Similaires

Towards Geographic Inclusion in the Evaluation of Text-to-Image ModelsBeyond Aesthetics: Cultural Competence in Text-to-Image ModelsDIG In: Evaluating Disparities in Image Generations with Indicators for Geographic DiversityGeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image ModelsText-to-Image Models and Their Representation of People from Different Nationalities Engaging in ActivitiesUse of ChatGPT to Explore Gender and Geographic Disparities in Scientific Peer Review (Preprint)

Towards Geographic Inclusion in the Evaluation of Text-to-Image Models

Rapid progress in text-to-image generative models coupled with their deployment for visual content c

Beyond Aesthetics: Cultural Competence in Text-to-Image Models

Text-to-Image (T2I) models are being increasingly adopted in diverse global communities where they c

DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity

The unprecedented photorealistic results achieved by recent text-to-image generative systems and the

GeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models

Text-to-image (T2I) models are rapidly gaining popularity, yet their outputs often lack geographical

Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities

This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, d

Use of ChatGPT to Explore Gender and Geographic Disparities in Scientific Peer Review (Preprint)

BACKGROUND In the realm of scientific research, peer review serves as a co