paper-with-me

Papers

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

2025-10-07 · Firoj Alam, Ali Ezzat Shahroor, Md. Arid Hasan, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Mohamed Bayan Kmainasi, Shammur Absar Chowdhury, Basel Mousi, Fahim Dalvi, Nadir Durrani, Natasa Milic-Frayling arxiv

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low-resource and underrepresented languages. We introduce OASIS, a large-scale culturally grounded multimodal QA dataset covering images, text, and speech. OASIS is built with EverydayMMQA, a scalable semi-automatic framework for creating localized spoken and visual QA resources, supported by multi-stage human-in-the-loop validation. OASIS contains approximately 0.92M real images and 14.8M QA pairs, including 3.7M spoken questions, with 383 hours of human-recorded speech, and 20K hours of voice-cloned speech, from 42 speakers. It supports four input settings: text-only, speech-only, text+image, and speech+image. The dataset focuses on English and Arabic varieties across 18 countries, covering Modern Standard Arabic (MSA) as well as dialectal Arabic. It is designed to evaluate models beyond object recognition, targeting pragmatic, commonsense, and culturally grounded reasoning in real-world scenarios. We benchmark four closed-source models, three open-source models, and one fine-tuned model on OASIS. The framework and dataset will be made publicly available to the community. https://huggingface.co/datasets/QCRI/OASIS

📄 PDF Abstract BibTeX arXiv:2510.06371

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringObject Recognition

Similar Papers 제목 키워드 기반

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

2025-08-07 · Yufei Gao, Jiaying Fei, Nuo Chen, Ruirui Chen 외 arxiv

Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings. We argue that this failure is not merely a lingu…

Machine Translation

Grounding Multilingual Multimodal LLMs With Cultural Knowledge

2025-08-10 · Jean de Dieu Nyandwi, Yueqi Song, Simran Khanuja, Graham Neubig arxiv

Multimodal Large Language Models excel in high-resource settings, but often misinterpret long-tail cultural entities and underperform in low-resource languages. To address this gap, we propose a data-centric approach tha…

Visual Question Answering

Toward Culturally Grounded Natural Language Processing

2026-03-27 · Sina Bagheri Nezhad arxiv

Multilingual NLP is often treated as a route to global inclusion, but linguistic coverage and cultural competence frequently diverge. This paper synthesizes over 50 papers spanning multilingual performance inequality, cr…

Cross-Lingual Transfer

IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs

2025-11-06 · Ali Faraz, Akash, Shaharukh Khan, Raja Kolla 외 arxiv

Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions about their performance in culturally diver…

Multimodal Machine TranslationVisual Question Answering

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture

2025-09-23 · Arijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka 외 arxiv

We introduce DRISHTIKON, a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture, designed to evaluate the cultural understanding of generative AI systems. Unlike existing benchma…