paper-with-me

Papers

VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding

2026-01-12 · Haorui Yu, Diji Yang, Hang He, Fengrui Zhang, Qiufeng Yi arxiv

We introduce VULCA-Bench, a multicultural art-critique benchmark for evaluating Vision-Language Models' (VLMs) cultural understanding beyond surface-level visual perception. Existing VLM benchmarks predominantly measure L1-L2 capabilities (object recognition, scene description, and factual question answering) while under-evaluate higher-order cultural interpretation. VULCA-Bench contains 7,410 matched image-critique pairs spanning eight cultural traditions, with Chinese-English bilingual coverage. We operationalise cultural understanding using a five-layer framework (L1-L5, from Visual Perception to Philosophical Aesthetics), instantiated as 225 culture-specific dimensions and supported by expert-written bilingual critiques. Our pilot results indicate that higher-layer reasoning (L3-L5) is consistently more challenging than visual and technical analysis (L1-L2). The dataset, evaluation scripts, and annotation tools are available under CC BY 4.0 at https://github.com/yha9806/VULCA-Bench.

📄 PDF Abstract BibTeX arXiv:2601.07986

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringObject Recognition

Similar Papers 제목 키워드 기반

WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

2024-10-16 · Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha 외

Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we …

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models

2024-06-28 · Mehar Bhatia, Sahithya Ravi, Aditya Chinchure, EunJeong Hwang 외

Despite recent advancements in vision-language models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to te…

DiversityRetrievalVisual Grounding

VULCAN: Vision-Language-Model Enhanced Multi-Agent Cooperative Navigation for Indoor Fire-Disaster Response

2026-04-14 · Shengding Liu, Qiben Yan arxiv

Indoor fire disasters pose severe challenges to autonomous search and rescue due to dense smoke, high temperatures, and dynamically evolving indoor environments. In such time-critical scenarios, multi-agent cooperative n…

M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks

2024-07-04 · Florian Schneider, Sunayana Sitaram

Since the release of ChatGPT, the field of Natural Language Processing has experienced rapid advancements, particularly in Large Language Models (LLMs) and their multimodal counterparts, Large Multimodal Models (LMMs). D…

Outlier Detection

Benchmarking Performance of Deep Learning Model for Material Segmentation on Two HPC Systems

2023-07-27 · Warren R. Williams, S. Ross Glandon, Luke L. Morris, Jing-Ru C. Cheng

Performance Benchmarking of HPC systems is an ongoing effort that seeks to provide information that will allow for increased performance and improve the job schedulers that manage these systems. We develop a benchmarking…

BenchmarkingGPUMaterial Segmentation