Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

Domaine:

natural language processinggeospatial

Type de record:

paperdataset
Créateur:
BöcNosPauIan
Hôte:avatar
Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and topological computation despite storing considerable geographic knowledge. Existing benchmarks localize these failures only partially: they are synthetic or smallscale, largely monolingual, and offer limited control over geographic coverage. We introduce MultiGlobeQA, a multilingual benchmark of 46,060 question-answer pairs spanning 14 spatial-function families and 15 answer formats, with execution-based ground truth over three knowledge graphs. It covers 201 countries and territories via income- and density-stratified sampling, with parallel questions in English and 16 additional high- and low-resource languages. Across parametric, reasoning, and agentic settings, LLMs collapse on tasks requiring grid indexing and shape computation, while topological relations and directions fare best. Retrieval and tool use yield considerable gains, yet performance plateaus below two thirds even when gold facts are supplied, indicating that computation, not access to knowledge, is the bottleneck. Models also underperform on low-income regions, a gap that gold facts widen rather than close.

Visit

arxiv.org

Tasks

question answering

Tags

Computation and LanguageArtificial IntelligenceInformation RetrievalI.2.7; H.3.3

Similaires

A Culturally-diverse Multilingual Multimodal Video Benchmark & ModelA Culturally-diverse Multilingual Multimodal Video Benchmark & ModelLearn Globally, Speak Locally: Bridging the Gaps in Multilingual ReasoningMacaron: A Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-FillingMacaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-FillingLinguini: A benchmark for language-agnostic linguistic reasoning

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understa

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understa

Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning

Large Language Models (LLMs) have achieved strong performance in domains like mathematics, factual q

Macaron: A Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling

Evaluating whether large language models (LLMs) genuinely understand the world from a non-English pe

Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling

Multilingual benchmarks rarely test reasoning over culturally grounded premises: translated datasets

Linguini: A benchmark for language-agnostic linguistic reasoning

We propose a new benchmark to measure a language model's linguistic reasoning skills without relying