paper-with-me

Papers

Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness

2025-09-27 · Yuchen Song, Andong Chen, Wenxin Zhu, Kehai Chen, Xuefeng Bai, Muyun Yang, Tiejun Zhao arxiv

Cultural awareness capabilities have emerged as a critical capability for Multimodal Large Language Models (MLLMs). However, current benchmarks lack progressed difficulty in their task design and are deficient in cross-lingual tasks. Moreover, current benchmarks often use real-world images. Each real-world image typically contains one culture, making these benchmarks relatively easy for MLLMs. Based on this, we propose C$^3$B (Comics Cross-Cultural Benchmark), a novel multicultural, multitask and multilingual cultural awareness capabilities benchmark. C$^3$B comprises over 2000 images and over 18000 QA pairs, constructed on three tasks with progressed difficulties, from basic visual recognition to higher-level cultural conflict understanding, and finally to cultural content generation. We conducted evaluations on 11 open-source MLLMs, revealing a significant performance gap between MLLMs and human performance. The gap demonstrates that C$^3$B poses substantial challenges for current MLLMs, encouraging future research to advance the cultural awareness capabilities of MLLMs.

📄 PDF Abstract BibTeX arXiv:2510.00041

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs

2025-05-16 · Pengju Xu, Yan Wang, Shuyuan Zhang, Xuan Zhou 외

Recent progress in Multimodal Large Language Models (MLLMs) have significantly enhanced the ability of artificial intelligence systems to understand and generate multimodal content. However, these models often exhibit li…

BenchmarkingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

2026-08-03 · Xianjing Han, Yuhan Su, Yang Deng, Dong Ma 외 arxiv

Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, …

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

2026-08-04 · Xiaolin Chen, Xuemeng Song, Wenhao Shi, Xianjing Han 외 arxiv

Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotion…

Explanation Generation

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

2025-10-17 · Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad 외 arxiv

Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularly for regions in the Middle East and Africa. In this paper, we introduce EgMM-Corpus, a multimodal dataset dedicated to …

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture

2025-09-23 · Arijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka 외 arxiv

We introduce DRISHTIKON, a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture, designed to evaluate the cultural understanding of generative AI systems. Unlike existing benchma…