Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

Domain:

natural language processing

Record type:

dataset
Creator:
AlaShaHasAli
Publisher:
arXiv
Host:avatar
Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low-resource and underrepresented languages. We introduce OASIS, a large-scale culturally grounded multimodal QA dataset covering images, text, and speech. OASIS is built with EverydayMMQA, a scalable semi-automatic framework for creating localized spoken and visual QA resources, supported by multi-stage human-in-the-loop validation. OASIS contains approximately 0.92M real images and 14.8M QA pairs, including 3.7M spoken questions, with 383 hours of human-recorded speech, and 20K hours of voice-cloned speech, from 42 speakers. It supports four input settings: text-only, speech-only, text+image, and speech+image. The dataset focuses on English and Arabic varieties across 18 countries, covering Modern Standard Arabic (MSA) as well as dialectal Arabic. It is designed to evaluate models beyond object recognition, targeting pragmatic, commonsense, and culturally grounded reasoning in real-world scenarios. We benchmark four closed-source models, three open-source models, and one fine-tuned model on OASIS. The framework and dataset will be made publicly available to the community. huggingface.co Multimodal Foundation Models, Large Language Models, Native, Multilingual, Language Diversity, Contextual Understanding, Culturally Informed

Visit

doi.orgarxiv.org

Tasks

question answering

Tags

Computation and Language (cs.CL)Artificial Intelligence (cs.AI)FOS: Computer and information sciencesFOS: Computer and information sciencesF.2.2; I.2.768T50

Licenses

Creative Commons Attribution Non Commercial Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-sa/4.0/legalcode

Similar

EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA

EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA),