paper-with-me

홈 › Papers

Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning

2025-09-27 · Haorui Yu, Yang Zhao, Yijia Chu, Qiufeng Yi arxiv

Vision-Language Models (VLMs) often appear culturally competent but rely on superficial pattern matching rather than genuine cultural understanding. We introduce a diagnostic framework to probe VLM reasoning on fire-themed cultural imagery through both classification and explanation analysis. Testing multiple models on Western festivals, non-Western traditions, and emergency scenes reveals systematic biases: models correctly identify prominent Western festivals but struggle with underrepresented cultural events, frequently offering vague labels or dangerously misclassifying emergencies as celebrations. These failures expose the risks of symbolic shortcuts and highlight the need for cultural evaluation beyond accuracy metrics to ensure interpretable and fair multimodal systems.

📄 PDF Abstract BibTeX arXiv:2509.23311

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting

2024-04-16 · Huihan Li, Liwei Jiang, Jena D. Hwang, Hyunwoo Kim 외

As the utilization of large language models (LLMs) has proliferated world-wide, it is crucial for them to have adequate knowledge and fair representation for diverse global cultures. In this work, we uncover culture perc…

DiversityFairness

Seeing Culture: A Benchmark for Visual Reasoning and Grounding

2025-09-20 · Burak Satar, Zhixin Ma, Patrick A. Irawan, Wilfried A. Mulyawan 외 arxiv

Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergen…

Visual Question AnsweringVisual Reasoning

Seeing What Is Not There: Learning Context to Determine Where Objects Are Missing

2017-02-26 · CVPR 2017 7 · Jin Sun, David W. Jacobs

Most of computer vision focuses on what is in an image. We propose to train a standalone object-centric context representation to perform the opposite task: seeing what is not there. Given an image, our context model can…

Objectobject-detectionObject Detection

Benchmarking Vision Language Models for Cultural Understanding

2024-07-15 · Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy 외

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed…

BenchmarkingQuestion AnsweringScene UnderstandingVisual Question Answering

Attributing Culture-Conditioned Generations to Pretraining Corpora

2024-12-30 · Huihan Li, Arnav Goel, Keyu He, Xiang Ren

In open-ended generative tasks like narrative writing or dialogue, large language models often exhibit cultural biases, showing limited knowledge and generating templated outputs for less prevalent cultures. Recent works…

Memorization