paper-with-me

홈 › Papers

World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models

2025-11-27 · Eunsu Kim, Junyeong Park, Na Min An, Junseong Kim, Hitesh Laxmichand Patel, Jiho Jin, Julia Kruk, Amit Agarwal, Srikant Panda, Fenal Ashokbhai Ilasariya, Hyunjung Shim, Alice Oh arxiv

In a globalized world, cultural elements from diverse origins frequently appear together within a single visual scene. We refer to these as culture mixing scenarios, yet how Large Vision-Language Models (LVLMs) perceive them remains underexplored. We investigate culture mixing as a critical challenge for LVLMs and examine how current models behave when cultural items from multiple regions appear together. To systematically analyze these behaviors, we construct CultureMix, a food Visual Question Answering (VQA) benchmark with 23k diffusion-generated, human-verified culture mixing images across four subtasks: (1) food-only, (2) food+food, (3) food+background, and (4) food+food+background. Evaluating 10 LVLMs, we find consistent failures to preserve individual cultural identities in mixed settings. Models show strong background reliance, with accuracy dropping 14% when cultural backgrounds are added to food-only baselines, and they produce inconsistent predictions for identical foods across different contexts. To address these limitations, we explore three robustness strategies. We find supervised fine-tuning using a diverse culture mixing dataset substantially improve model consistency and reduce background sensitivity. We call for increased attention to culture mixing scenarios as a critical step toward developing LVLMs capable of operating reliably in culturally diverse real-world environments.

📄 PDF Abstract BibTeX arXiv:2511.22787

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Shedding light on social learning

2023-10-13 · Kingsley J. A. Cox, Paul R. Adams

Culture involves the origination and transmission of ideas, but the conditions in which culture can emerge and evolve are unclear. We constructed and studied a highly simplified neural-network model of these processes. I…

Unmixing Diffusion for Self-Supervised Hyperspectral Image Denoising

2024-01-01 · CVPR 2024 1 · Haijin Zeng, JieZhang Cao, Kai Zhang, Yongyong Chen 외

Hyperspectral images (HSIs) have extensive applications in various fields such as medicine agriculture and industry. Nevertheless acquiring high signal-to-noise ratio HSI poses a challenge due to narrow-band spectral…

DenoisingHyperspectral Image DenoisingImage Denoising

Understanding Script-Mixing: A Case Study of Hindi-English Bilingual Twitter Users

2020-05-01 · LREC 2020 5 · Abhishek Srivastava, Kalika Bali, Monojit Choudhury

In a multi-lingual and multi-script society such as India, many users resort to code-mixing while typing on social media. While code-mixing has received a lot of attention in the past few years, it has mostly been studie…

Sentence

Neuron-Level Analysis of Cultural Understanding in Large Language Models

2025-10-09 · Taisei Yamamoto, Ryoma Kumon, Danushka Bollegala, Hitomi Yanaka arxiv

As large language models (LLMs) are increasingly deployed worldwide, ensuring their fair and comprehensive cultural understanding is important. However, LLMs exhibit cultural bias and limited awareness of underrepresente…

Natural Language Understanding

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

2026-08-04 · Xiaolin Chen, Xuemeng Song, Wenhao Shi, Xianjing Han 외 arxiv

Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotion…

Explanation Generation