Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Constructing Multilingual Visual-Text Datasets Revealing Visual Multilingual Ability of Vision Language Models

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
AtuAliHirKam
Hôte:avatar
Large language models (LLMs) have increased interest in vision language models (VLMs), which process image-text pairs as input. Studies investigating the visual understanding ability of VLMs have been proposed, but such studies are still preliminary because existing datasets do not permit a comprehensive evaluation of the fine-grained visual linguistic abilities of VLMs across multiple languages. To further explore the strengths of VLMs, such as GPT-4V \cite{openai2023GPT4}, we developed new datasets for the systematic and qualitative analysis of VLMs. Our contribution is four-fold: 1) we introduced nine vision-and-language (VL) tasks (including object recognition, image-text matching, and more) and constructed multilingual visual-text datasets in four languages: English, Japanese, Swahili, and Urdu through utilizing templates containing \textit{questions} and prompting GPT4-V to generate the \textit{answers} and the \textit{rationales}, 2) introduced a new VL task named \textit{unrelatedness}, 3) introduced rationales to enable human understanding of the VLM reasoning process, and 4) employed human evaluation to measure the suitability of proposed datasets for VL tasks. We show that VLMs can be fine-tuned on our datasets. Our work is the first to conduct such analyses in Swahili and Urdu. Also, it introduces \textit{rationales} in VL analysis, which played a vital role in the evaluation.

Visit

arxiv.org

Tasks

computer visionimage-text retrieval

Languages

Swahili

Tags

Computation and LanguageComputer Vision and Pattern Recognition

Similaires

Multilingual Diversity Improves Vision-Language RepresentationsVision-Based Multilingual Sign Language TranslationTHE EMERGENCE OF PERCEPTUAL GROUPING ABILITY IN NEURAL NETWORK MODELS OF VISUAL CORTEX OPERATIONMultilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language ModelsEverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QAMultilingual Practices, Critical Literacies, and Visual Culture: A Focus on African Contexts

Multilingual Diversity Improves Vision-Language Representations

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learnin

Vision-Based Multilingual Sign Language Translation

THE EMERGENCE OF PERCEPTUAL GROUPING ABILITY IN NEURAL NETWORK MODELS OF VISUAL CORTEX OPERATION

Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models

Recently, it has been found that monolingual English language models can be used as knowledge bases. Instead of structural knowledge base queries, masked sentences such as "Paris is the capital of [MASK]" are used as probes. We translate the established benchmarks

EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA),

Multilingual Practices, Critical Literacies, and Visual Culture: A Focus on African Contexts

Abstract In this essay, we review and comment on three books that focus on language, literacy, and