paper-with-me

홈 › Papers

VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering

2025-11-12 · Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le arxiv

Contemporary Visual Question Answering (VQA) systems remain constrained when confronted with culturally specific content, largely because cultural knowledge is under-represented in training corpora and the reasoning process is not rendered interpretable to end users. This paper introduces VietMEAgent, a multimodal explainable framework engineered for Vietnamese cultural understanding. The method integrates a cultural object detection backbone with a structured program generation layer, yielding a pipeline in which answer prediction and explanation are tightly coupled. A curated knowledge base of Vietnamese cultural entities serves as an explicit source of background information, while a dual-modality explanation module combines attention-based visual evidence with structured, human-readable textual rationales. We further construct a Vietnamese Cultural VQA dataset sourced from public repositories and use it to demonstrate the practicality of programming-based methodologies for cultural AI. The resulting system provides transparent explanations that disclose both the computational rationale and the underlying cultural context, supporting education and cultural preservation with an emphasis on interpretability and cultural sensitivity.

📄 PDF Abstract BibTeX arXiv:2511.09058

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringObject Detection

Similar Papers 제목 키워드 기반

Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications

2025-06-22 · Bushra Asseri, Estabraq Abdelaziz, Maha Al Mogren, Tayef Alhefdhi 외

Emotion recognition capabilities in multimodal AI systems are crucial for developing culturally responsive educational technologies, yet remain underexplored for Arabic language contexts where culturally appropriate lear…

Emotion Recognition

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

2025-10-17 · Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad 외 arxiv

Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularly for regions in the Middle East and Africa. In this paper, we introduce EgMM-Corpus, a multimodal dataset dedicated to …

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture

2025-09-23 · Arijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka 외 arxiv

We introduce DRISHTIKON, a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture, designed to evaluate the cultural understanding of generative AI systems. Unlike existing benchma…

Creating a Lens of Chinese Culture: A Multimodal Dataset for Chinese Pun Rebus Art Understanding

2024-06-14 · Tuo Zhang, Tiantian Feng, Yibin Ni, Mengqin Cao 외

Large vision-language models (VLMs) have demonstrated remarkable abilities in understanding everyday content. However, their performance in the domain of art, particularly culturally rich art forms, remains less explored…

In-Context Learning

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

2026-08-31 · Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir 외 arxiv

We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination …

Visual Question AnsweringText-to-Image Generation