paper-with-me

홈 › Papers

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images

2026-04-08 · Yuechen Jiang, Enze Zhang, Md Mohsinul Kabir, Qianqian Xie, Stavroula Golfomitsou, Konstantinos Arvanitis, Sophia Ananiadou arxiv

Recent advances in vision-language models (VLMs) have improved image captioning for cultural heritage. However, inferring structured cultural metadata (e.g., creator, origin, period) from visual input remains underexplored. We introduce a multi-category, cross-cultural benchmark for this task and evaluate VLMs using an LLM-as-Judge framework that measures semantic alignment with reference annotations. To assess cultural reasoning, we report exact-match, partial-match, and attribute-level accuracy across cultural regions. Results show that models capture fragmented signals and exhibit substantial performance variation across cultures and metadata types, leading to inconsistent and weakly grounded predictions. These findings highlight the limitations of current VLMs in structured cultural metadata inference beyond visual perception.

📄 PDF Abstract BibTeX arXiv:2604.07338

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

Supervised Semantic Differential for Cross-Cultural Concept Analysis: A Case Study of Human Affect

2026-05-27 · Jan Sikora, Paweł Lenartowicz, Hubert Plisiecki arxiv

Cross-cultural comparison of psychological meaning requires methods that go beyond word-level translation and examine how semantic dimensions are organized across languages. We introduce a cross-lingual extension of the …

Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts

2025-02-20 · Sara Ghaboura, Ketan More, Ritesh Thawkar, Wafa Alghallabi 외

Understanding historical and cultural artifacts demands human expertise and advanced computational techniques, yet the process remains complex and time-intensive. While large multimodal models offer promising support, th…

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

2026-03-22 · Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Shubhashis Roy Dipta, Rubaya Tabassum 외 arxiv

Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVers…

Visual Question AnsweringVisual Grounding

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

2026-08-13 · Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi 외 arxiv

Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks r…

Speaker Verification

From Words to Worlds: Benchmarking Cross-Cultural Cultural Understanding in Machine Translation

2026-03-18 · Bangju Han, Yingqi Wang, Huang Qing, Tiyuan Li 외 arxiv

Culture-expressions, such as idioms, slang, and culture-specific items (CSIs), are pervasive in natural language and encode meanings that go beyond literal linguistic form. Accurately translating such expressions remains…

Machine Translation