Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AlaShaHasAli
Éditeur:
arXiv
Hôte:avatar
Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low-resource and underrepresented languages. We introduce OASIS, a large-scale culturally grounded multimodal QA dataset covering images, text, and speech. OASIS is built with EverydayMMQA, a scalable semi-automatic framework for creating localized spoken and visual QA resources, supported by multi-stage human-in-the-loop validation. OASIS contains approximately 0.92M real images and 14.8M QA pairs, including 3.7M spoken questions, with 383 hours of human-recorded speech, and 20K hours of voice-cloned speech, from 42 speakers. It supports four input settings: text-only, speech-only, text+image, and speech+image. The dataset focuses on English and Arabic varieties across 18 countries, covering Modern Standard Arabic (MSA) as well as dialectal Arabic. It is designed to evaluate models beyond object recognition, targeting pragmatic, commonsense, and culturally grounded reasoning in real-world scenarios. We benchmark four closed-source models, three open-source models, and one fine-tuned model on OASIS. The framework and dataset will be made publicly available to the community. huggingface.co Multimodal Foundation Models, Large Language Models, Native, Multilingual, Language Diversity, Contextual Understanding, Culturally Informed

Visit

doi.orgarxiv.org

Tasks

question answering

Tags

Computation and Language (cs.CL)Artificial Intelligence (cs.AI)FOS: Computer and information sciencesFOS: Computer and information sciencesF.2.2; I.2.768T50

Licenses

Creative Commons Attribution Non Commercial Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-sa/4.0/legalcode

Similaires

EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA

EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA),