paper-with-me

Papers

Vision-Language Models under Cultural and Inclusive Considerations

2024-07-08 · Antonia Karamolegkou, Phillip Rust, Yong Cao, Ruixiang Cui, Anders Søgaard, Daniel Hershcovich

Large vision-language models (VLMs) can assist visually impaired people by describing images from their daily lives. Current evaluation datasets may not reflect diverse cultural user backgrounds or the situational context of this use case. To address this problem, we create a survey to determine caption preferences and propose a culture-centric evaluation benchmark by filtering VizWiz, an existing dataset with images taken by people who are blind. We then evaluate several VLMs, investigating their reliability as visual assistants in a culturally diverse setting. While our results for state-of-the-art models are promising, we identify challenges such as hallucination and misalignment of automatic evaluation metrics with human judgment. We make our survey, data, code, and model outputs publicly available.

📄 PDF Abstract BibTeX arXiv:2407.06177

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationSurvey

Similar Papers 제목 키워드 기반

Culture is Everywhere: A Call for Intentionally Cultural Evaluation

2025-09-01 · Juhyun Oh, Inha Cha, Michael Saxon, Hyunseung Lim 외 arxiv

The prevailing ``trivia-centered paradigm'' for evaluating the cultural alignment of large language models (LLMs) is increasingly inadequate as these models become more advanced and widely deployed. Existing approaches t…

No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models

2024-05-22 · Angéline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang 외

We study cultural and socioeconomic diversity in contrastive vision-language models (VLMs). Using a broad range of benchmark datasets and evaluation metrics, we bring to attention several important findings. First, the c…

Diversitygeo-localization

JEEM: Vision-Language Understanding in Four Arabic Dialects

2025-03-27 · Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali, Abdelrahman Mohamed 외

We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. JEEM includes the tasks of image …

Image CaptioningQuestion AnsweringVisual Question Answering

Rice-VL: Evaluating Vision-Language Models for Cultural Understanding Across ASEAN Countries

2025-12-01 · Tushar Pranav, Eshan Pandey, Austria Lyka Diane Bala, Aman Chadha 외 arxiv

Vision-Language Models (VLMs) excel in multimodal tasks but often exhibit Western-centric biases, limiting their effectiveness in culturally diverse regions like Southeast Asia (SEA). To address this, we introduce RICE-V…

Visual Question AnsweringVisual Grounding

IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs

2025-11-06 · Ali Faraz, Akash, Shaharukh Khan, Raja Kolla 외 arxiv

Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions about their performance in culturally diver…

Multimodal Machine TranslationVisual Question Answering