paper-with-me

Papers

How Culturally Aware are Vision-Language Models?

2024-05-24 · Olena Burda-Lassen, Aman Chadha, Shashank Goswami, Vinija Jain

An image is often said to be worth a thousand words, and certain images can tell rich and insightful stories. Can these stories be told via image captioning? Images from folklore genres, such as mythology, folk dance, cultural signs, and symbols, are vital to every culture. Our research compares the performance of four popular vision-language models (GPT-4V, Gemini Pro Vision, LLaVA, and OpenFlamingo) in identifying culturally specific information in such images and creating accurate and culturally sensitive image captions. We also propose a new evaluation metric, Cultural Awareness Score (CAS), dedicated to measuring the degree of cultural awareness in image captions. We provide a dataset MOSAIC-1.5k, labeled with ground truth for images containing cultural background and context, as well as a labeled dataset with assigned Cultural Awareness Scores that can be used with unseen data. Creating culturally appropriate image captions is valuable for scientific research and can be beneficial for many practical applications. We envision that our work will promote a deeper integration of cultural sensitivity in AI applications worldwide. By making the dataset and Cultural Awareness Score available to the public, we aim to facilitate further research in this area, encouraging the development of more culturally aware AI systems that respect and celebrate global diversity.

📄 PDF Abstract BibTeX arXiv:2405.17475

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

CIC: A Framework for Culturally-Aware Image Captioning

2024-02-08 · Youngsik Yun, Jihie Kim

Image Captioning generates descriptive sentences from images using Vision-Language Pre-trained models (VLPs) such as BLIP, which has improved greatly. However, current methods lack the generation of detailed descriptive …

DescriptiveImage CaptioningQuestion AnsweringVisual Question Answering+1

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

2025-10-17 · Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad 외 arxiv

Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularly for regions in the Middle East and Africa. In this paper, we introduce EgMM-Corpus, a multimodal dataset dedicated to …

Toward Culturally Grounded Natural Language Processing

2026-03-27 · Sina Bagheri Nezhad arxiv

Multilingual NLP is often treated as a route to global inclusion, but linguistic coverage and cultural competence frequently diverge. This paper synthesizes over 50 papers spanning multilingual performance inequality, cr…

Cross-Lingual Transfer

Culturally Aware and Adapted NLP: A Taxonomy and a Survey of the State of the Art

2024-06-06 · Chen Cecilia Liu, Iryna Gurevych, Anna Korhonen

The surge of interest in culturally aware and adapted Natural Language Processing (NLP) has inspired much recent research. However, the lack of common understanding of the concept of "culture" has made it difficult to ev…

Multilingual Vision-Language Models, A Survey

2025-09-26 · Andrei-Alexandru Manea, Jindřich Libovický arxiv

This survey examines multilingual vision-language models that process text and images across languages. We review 33 models and 23 benchmarks, spanning encoder-only and generative architectures, and identify a key tensio…

Contrastive Learning