Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Beyond Aesthetics: Cultural Competence in Text-to-Image Models

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
KanAhmAndPra
Hôte:avatar
Text-to-Image (T2I) models are being increasingly adopted in diverse global communities where they create visual representations of their unique cultures. Current T2I benchmarks primarily focus on faithfulness, aesthetics, and realism of generated images, overlooking the critical dimension of cultural competence. In this work, we introduce a framework to evaluate cultural competence of T2I models along two crucial dimensions: cultural awareness and cultural diversity, and present a scalable approach using a combination of structured knowledge bases and large language models to build a large dataset of cultural artifacts to enable this evaluation. In particular, we apply this approach to build CUBE (CUltural BEnchmark for Text-to-Image models), a first-of-its-kind benchmark to evaluate cultural competence of T2I models. CUBE covers cultural artifacts associated with 8 countries across different geo-cultural regions and along 3 concepts: cuisine, landmarks, and art. CUBE consists of 1) CUBE-1K, a set of high-quality prompts that enable the evaluation of cultural awareness, and 2) CUBE-CSpace, a larger dataset of cultural artifacts that serves as grounding to evaluate cultural diversity. We also introduce cultural diversity as a novel T2I evaluation component, leveraging quality-weighted Vendi score. Our evaluations reveal significant gaps in the cultural awareness of existing models across countries and provide valuable insights into the cultural diversity of T2I outputs for under-specified prompts. Our methodology is extendable to other cultural regions and concepts, and can facilitate the development of T2I models that better cater to the global population. NeurIPS 2024 camera-ready version

Visit

arxiv.org

Tasks

computer visionimage-text retrieval

Tags

Computer Vision and Pattern Recognition

Similaires

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language ModelsDecomposed evaluations of geographic disparities in text-to-image modelsGeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image ModelsTowards Geographic Inclusion in the Evaluation of Text-to-Image ModelsText-to-Image Models and Their Representation of People from Different Nationalities Engaging in ActivitiesBeyond Models: A Framework for Contextual and Cultural Intelligence in African AI Deployment

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks

Decomposed evaluations of geographic disparities in text-to-image models

Recent work has identified substantial disparities in generated images of different geographic regio

GeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models

Text-to-image (T2I) models are rapidly gaining popularity, yet their outputs often lack geographical

Towards Geographic Inclusion in the Evaluation of Text-to-Image Models

Rapid progress in text-to-image generative models coupled with their deployment for visual content c

Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities

This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, d

Beyond Models: A Framework for Contextual and Cultural Intelligence in African AI Deployment

While global AI development prioritizes model performance and computational scale, meaningful deploy