Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
YueZhaCheHou, Peng
Hôte:avatar
Multilingual document and scene text understanding plays an important role in applications such as search, finance, and public services. However, most existing benchmarks focus on high-resource languages and fail to evaluate models in realistic multilingual environments. In Southeast Asia, the diversity of languages, complex writing systems, and highly varied document types make this challenge even greater. We introduce SEA-Vision, a benchmark that jointly evaluates Document Parsing and Text-Centric Visual Question Answering (TEC-VQA) across 11 Southeast Asian languages. SEA-Vision contains 15,234 document parsing pages from nine representative document types, annotated with hierarchical page-, block-, and line-level labels. It also provides 7,496 TEC-VQA question-answer pairs that probe text recognition, numerical calculation, comparative analysis, logical reasoning, and spatial understanding. To make such multilingual, multi-task annotation feasible, we design a hybrid pipeline for Document Parsing and TEC-VQA. It combines automated filtering and scoring with MLLM-assisted labeling and lightweight native-speaker verification, greatly reducing manual labeling while maintaining high quality. We evaluate several leading multimodal models and observe pronounced performance degradation on low-resource Southeast Asian languages, highlighting substantial remaining gaps in multilingual document and scene text understanding. We believe SEA-Vision will help drive global progress in document and scene text understanding. Accepted By CVPR2026

Visit

arxiv.org

Tasks

question answering

Tags

Computation and Language

Similaires

Comprehensive Benchmark Datasets for Amharic Scene Text Detection and RecognitionKhmerST: A Low-Resource Khmer Scene Text Detection and Recognition BenchmarkPragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource SettingsCross-Lingual Learning in Multilingual Scene Text RecognitionFleurs-SLU: A Massively Multilingual Benchmark for Spoken Language UnderstandingEnhanced Arabic Document Understanding Using Vision-Language Models

Comprehensive Benchmark Datasets for Amharic Scene Text Detection and Recognition

Ethiopic/Amharic script is one of the oldest African writing systems, which serves at least 23 langu

KhmerST: A Low-Resource Khmer Scene Text Detection and Recognition Benchmark

Developing effective scene text detection and recognition models hinges on extensive training data,

PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings

India's 22 official languages create a critical accessibility barrier: the majority of medical docum

Cross-Lingual Learning in Multilingual Scene Text Recognition

In this paper, we investigate cross-lingual learning (CLL) for multilingual scene text recognition (

Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding

Spoken language understanding (SLU) is indispensable for half of all living languages that lack a fo

Enhanced Arabic Document Understanding Using Vision-Language Models