Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TDBench: A Benchmark for Top-Down Image Understanding with Reliability Analysis of Vision-Language Models

Record type:

paperdataset
Creator:
HouZhaXu,Fan
Host:avatar
Top-down images play an important role in safety-critical settings such as autonomous navigation and aerial surveillance, where they provide holistic spatial information that front-view images cannot capture. Despite this, Vision Language Models (VLMs) are mostly trained and evaluated on front-view benchmarks, leaving their performance in the top-down setting poorly understood. Existing evaluations also overlook a unique property of top-down images: their physical meaning is preserved under rotation. In addition, conventional accuracy metrics can be misleading, since they are often inflated by hallucinations or "lucky guesses", which obscures a model's true reliability and its grounding in visual evidence. To address these issues, we introduce TDBench, a benchmark for top-down image understanding that includes 2000 curated questions for each rotation. We further propose RotationalEval (RE), which measures whether models provide consistent answers across four rotated views of the same scene, and we develop a reliability framework that separates genuine knowledge from chance. Finally, we conduct four case studies targeting underexplored real-world challenges. By combining rigorous evaluation with reliability metrics, TDBench not only benchmarks VLMs in top-down perception but also provides a new perspective on trustworthiness, guiding the development of more robust and grounded AI systems. Project homepage: github.com

Visit

arxiv.org

Tasks

computer vision

Tags

Machine LearningArtificial IntelligenceComputation and Language

Similar

Benchmarking Vision Language Models for Cultural UnderstandingEnhanced Arabic Document Understanding Using Vision-Language ModelsMangaUB: A Manga Understanding Benchmark for Large Multimodal ModelsLet's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of SportsCultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 CountriesSEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia

Benchmarking Vision Language Models for Cultural Understanding

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLM

Enhanced Arabic Document Understanding Using Vision-Language Models

MangaUB: A Manga Understanding Benchmark for Large Multimodal Models

Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panel

Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries

Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understa

SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia

Multilingual document and scene text understanding plays an important role in applications such as s