paper-with-me

Papers

Cultural Evaluations of Vision-Language Models Have a Lot to Learn from Cultural Theory

2025-05-28 · Srishti Yadav, Lauren Tilton, Maria Antoniak, Taylor Arnold, Jiaang Li, Siddhesh Milind Pawar, Antonia Karamolegkou, Stella Frank, Zhaochong An, Negar Rostamzadeh, Daniel Hershcovich, Serge Belongie, Ekaterina Shutova

Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in understanding how they encode cultural nuances. While individual aspects of this problem have been studied, we still lack a comprehensive framework for systematically identifying and annotating the nuanced cultural dimensions present in images for VLMs. This position paper argues that foundational methodologies from visual culture studies (cultural studies, semiotics, and visual studies) are necessary for cultural analysis of images. Building upon this review, we propose a set of five frameworks, corresponding to cultural dimensions, that must be considered for a more complete analysis of the cultural competencies of VLMs.

📄 PDF Abstract BibTeX arXiv:2505.22793

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityPosition

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Vision Language Models are Confused Tourists

2025-11-21 · Patrick Amadeus Irawan, Ikhlasul Akmal Hanif, Muhammad Dehan Al Kautsar, Genta Indra Winata 외 arxiv

Although the cultural dimension has been one of the key aspects in evaluating Vision-Language Models (VLMs), their ability to remain stable across diverse cultural inputs remains largely untested, despite being crucial t…

Adversarial Robustness

Navigating Cultural Chasms: Exploring and Unlocking the Cultural POV of Text-To-Image Models

2023-10-03 · Mor Ventura, Eyal Ben-David, Anna Korhonen, Roi Reichart

Text-To-Image (TTI) models, such as DALL-E and StableDiffusion, have demonstrated remarkable prompt-based image generation capabilities. Multilingual encoders may have a substantial impact on the cultural agency of these…

Image GenerationVisual Question Answering (VQA)

Rice-VL: Evaluating Vision-Language Models for Cultural Understanding Across ASEAN Countries

2025-12-01 · Tushar Pranav, Eshan Pandey, Austria Lyka Diane Bala, Aman Chadha 외 arxiv

Vision-Language Models (VLMs) excel in multimodal tasks but often exhibit Western-centric biases, limiting their effectiveness in culturally diverse regions like Southeast Asia (SEA). To address this, we introduce RICE-V…

Visual Question AnsweringVisual Grounding

Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses

2026-01-19 · Chongyuan Dai, Yaling Shen, Jinpeng Hu, Zihan Gao 외 arxiv

Culture serves as a fundamental determinant of human affective processing and profoundly shapes how individuals perceive and interpret emotional stimuli. Despite this intrinsic link extant evaluations regarding cultural …

On the Cultural Anachronism and Temporal Reasoning in Vision Language Models

2026-05-14 · Mukul Ranjan, Prince Jha, Khushboo Kumari, Zhiqiang Shen arxiv

Vision-Language Models (VLMs) are increasingly applied to cultural heritage materials, from digital archives to educational platforms. This work identifies a fundamental issue in how these models interpret historical art…