paper-with-me

Papers

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries

2025-01-02 · Shudong Liu, Yiqiao Jin, Cheng Li, Derek F. Wong, Qingsong Wen, Lichao Sun, Haipeng Chen, Xing Xie, Jindong Wang

Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understanding, often misinterpreting symbols, gestures, and artifacts due to biases in predominantly Western-centric training data. In this paper, we construct CultureVerse, a large-scale multimodal benchmark covering 19, 682 cultural concepts, 188 countries/regions, 15 cultural concepts, and 3 question types, with the aim of characterizing and improving VLMs' multicultural understanding capabilities. Then, we propose CultureVLM, a series of VLMs fine-tuned on our dataset to achieve significant performance improvement in cultural understanding. Our evaluation of 16 models reveals significant disparities, with a stronger performance in Western concepts and weaker results in African and Asian contexts. Fine-tuning on our CultureVerse enhances cultural perception, demonstrating cross-cultural, cross-continent, and cross-dataset generalization without sacrificing performance on models' general VLM benchmarks. We further present insights on cultural generalization and forgetting. We hope that this work could lay the foundation for more equitable and culturally aware multimodal AI systems.

📄 PDF Abstract BibTeX arXiv:2501.01282

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Benchmarking Vision Language Models for Cultural Understanding

2024-07-15 · Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy 외

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed…

BenchmarkingQuestion AnsweringScene UnderstandingVisual Question Answering

From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models

2024-06-28 · Mehar Bhatia, Sahithya Ravi, Aditya Chinchure, EunJeong Hwang 외

Despite recent advancements in vision-language models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to te…

DiversityRetrievalVisual Grounding

VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding

2026-01-12 · Haorui Yu, Diji Yang, Hang He, Fengrui Zhang 외 arxiv

We introduce VULCA-Bench, a multicultural art-critique benchmark for evaluating Vision-Language Models' (VLMs) cultural understanding beyond surface-level visual perception. Existing VLM benchmarks predominantly measure …

Question AnsweringObject Recognition

Cultural Evaluations of Vision-Language Models Have a Lot to Learn from Cultural Theory

2025-05-28 · Srishti Yadav, Lauren Tilton, Maria Antoniak, Taylor Arnold 외

Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in understanding how they encode cultural …

DiversityPosition

Rice-VL: Evaluating Vision-Language Models for Cultural Understanding Across ASEAN Countries

2025-12-01 · Tushar Pranav, Eshan Pandey, Austria Lyka Diane Bala, Aman Chadha 외 arxiv

Vision-Language Models (VLMs) excel in multimodal tasks but often exhibit Western-centric biases, limiting their effectiveness in culturally diverse regions like Southeast Asia (SEA). To address this, we introduce RICE-V…

Visual Question AnsweringVisual Grounding