Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning

Domaine:

natural language processing

Type de record:

paper
Commonsense is defined as the knowledge on which everyone agrees. However, certain types of commonsense knowledge are correlated with culture and geographic locations and they are only shared locally. For example, the scenes of wedding ceremonies vary across regions due to different customs influenced by historical and religious factors. Such regional characteristics, however, are generally omitted in prior work. In this paper, we construct a Geo-Diverse Visual Commonsense Reasoning dataset (GD-VCR) to test vision-and-language models{'} ability to understand cultural and geo-location-specific commonsense. In particular, we study two state-of-the-art Vision-and-Language models, VisualBERT and ViLBERT trained on VCR, a standard benchmark with images primarily from Western regions. We then evaluate how well the trained models can generalize to answering the questions in GD-VCR. We find that the performance of both models for non-Western regions including East Asia, South Asia, and Africa is significantly lower than that for Western region. We analyze the reasons behind the performance disparity and find that the performance gap is larger on QA pairs that: 1) are concerned with culture-related scenarios, e.g., weddings, religious activities, and festivals; 2) require high-level geo-diverse commonsense reasoning rather than low-order perception and recognition. Dataset and code are released at GD-VCR.

Visit

aclanthology.orgwww.youtube.com

Connected records

project

Tasks

commonsense reasoning

Tags

aclGD-VCR

Similaires

A Deep Learning-Based Bengali Visual Commonsense Reasoning SystemGeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsCommonsense Reasoning in Arab CultureXCOPA: A Multilingual Dataset for Causal Commonsense ReasoningGlobal PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and CulturesCross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World

A Deep Learning-Based Bengali Visual Commonsense Reasoning System

GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models

Recent work has shown that Pre-trained Language Models (PLMs) store the relational knowledge learned

Commonsense Reasoning in Arab Culture

Despite progress in Arabic large language models, such as Jais and AceGPT, their evaluation on commo

XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning

In order to simulate human language capacity, natural language processing systems must be able to re

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (

Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World

Large language models (LLMs) often reflect Western-centric biases, limiting their effectiveness in d