Logo Lanfrica

Multilingual Document Understanding: From Global Benchmarking to Arabic-Centric Evaluation

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
Hea
Éditeur:
Ayobin
Éditeur:
MBZ
Hôte:avatar
How well do modern document understanding systems truly comprehend the world’s written languages? Despite remarkable progress in vision-language models (VLMs), the ability to parse, recognize, and extract structured content from documents remains heavily skewed toward English and a handful of high-resource scripts. This thesis addresses the multilingual document understanding gap through two complementary contributions: a large-scale cross-lingual framework that exposes systemic failures across diverse writing systems, and a focused benchmark that dissects the unique challenges of Arabic, one of the most widely spoken yet computationally underserved languages. We first introduce DocAtlas, a framework for constructing high-fidelity OCR datasets and benchmarks covering 82 languages and 9 evaluation tasks through model-free annotation pipelines, yielding a 360K-page training corpus and a difficulty-stratified benchmark of 5,862 pages. Evaluating 16 state-of-the-art models reveals that low-resource scripts suffer 40–60% accuracy drops compared to high-resource counterparts, and that structured extraction plateaus at 71-73% TEDS regardless of language. We further demonstrate that Direct Preference Optimization (DPO) with rendering-derived ground truth achieves stable cross-lingual transfer, improving both in-domain (+1.9%) and out-of-domain (+1.8%) accuracy, where supervised fine-tuning degrades out-of-domain performance by up to 21%. Motivated by the persistent Arabic underperformance exposed by DocAtlas, we present KITAB-Bench, a comprehensive Arabic OCR benchmark spanning 9 domains and 36 sub-domains with 8,809 samples, covering tasks from basic text recognition to table extraction, chart understanding, diagram parsing, and end-to-end PDF-to-Markdown conversion. We introduce three evaluation metrics tailored to Arabic documents: MARS, CharTeX, and CODM. Our evaluation shows that modern VLMs outperform classical OCR by an average of 60% in character error rate, yet the best model achieves only 65% on PDF-to-Markdown conversion, exposing critical limitations in Arabic document understanding. Together, DocAtlas and KITAB-Bench provide a telescope-to-microscope view of multilingual document understanding: the former maps the global landscape, the latter dissects one of its most challenging scripts.