Logo Lanfrica

AI-Assisted Digitalization of Heterogeneous Government Archives in Low-Resource Countries: A Vision Language Model Case Study from Chad

Domaine:

digital infrastructure

Type de record:

paper
Créateur:
Chr
Éditeur:
Elsevier BV
Hôte:
The accumulation of millions of physical documents over decades has driven African governments to adopt digital technologies to improve document accessibility and reduce the risk of loss from unforeseen events. The specific linguistic, physical, institutional, and administrative characteristics of these collections have received almost no attention in the document AI literature. This paper addresses that gap through a case study of bilingual French-Arabic government document digitalisation in Chad, describing the design, implementation, and evaluation of an end-to-end AI pipeline for structured data extraction from a collection of approximately 50,000 administrative documents. We evaluate three open-source OCR tools (Tesseract, EasyOCR, and Surya) and demonstrate that all three are insufficient for structured field extraction on this document type, achieving mean confidence scores between 0.416 and 0.798 while producing no structured output, and that standard preprocessing reduces rather than improves performance. We show that Qwen2.5-VL 7B, a document-trained vision language model deployed in zero-shot configuration via Ollama on consumer hardware, achieves 70-78% validated field-level accuracy on humanannotated ground truth across high-confidence extractions, with 88-91% of documents classified as high confidence by the model's self-assessment across a 702-document operational batch, compared to 12% for the general-purpose LLaVA 7B, at 17.6 seconds per document with no internet connectivity. Ground truth validation on 28 annotated documents identifies four characterisable VLM-specific error patterns (organisation name hallucination on damaged letterheads, date field ambiguity on multi-date documents, handwritten document type misclassification, and reference number completion on partially legible fields) that are distinct from OCR failure modes and require dedicated post-extraction validation steps.

Similaires