paper-with-me

홈 › Papers

Creating a Lens of Chinese Culture: A Multimodal Dataset for Chinese Pun Rebus Art Understanding

2024-06-14 · Tuo Zhang, Tiantian Feng, Yibin Ni, Mengqin Cao, Ruying Liu, Katharine Butler, Yanjun Weng, Mi Zhang, Shrikanth S. Narayanan, Salman Avestimehr

Large vision-language models (VLMs) have demonstrated remarkable abilities in understanding everyday content. However, their performance in the domain of art, particularly culturally rich art forms, remains less explored. As a pearl of human wisdom and creativity, art encapsulates complex cultural narratives and symbolism. In this paper, we offer the Pun Rebus Art Dataset, a multimodal dataset for art understanding deeply rooted in traditional Chinese culture. We focus on three primary tasks: identifying salient visual elements, matching elements with their symbolic meanings, and explanations for the conveyed messages. Our evaluation reveals that state-of-the-art VLMs struggle with these tasks, often providing biased and hallucinated explanations and showing limited improvement through in-context learning. By releasing the Pun Rebus Art Dataset, we aim to facilitate the development of VLMs that can better understand and interpret culturally specific content, promoting greater inclusiveness beyond English-based corpora.

📄 PDF Abstract BibTeX arXiv:2406.10318

Code (1)

zhang-tuo-pdf/Pun-Rebus-Art-Benchmark 공식 구현

Tasks

In-Context Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Is Human Culture Locked by Evolution?

2023-10-28 · Hao Wang

Human culture has evolved for thousands of years and thrived in the era of Internet. Due to the availability of big data, we could do research on human culture by analyzing its representation such as user item rating val…

Recommendation Systems

TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs

2025-05-16 · Pengju Xu, Yan Wang, Shuyuan Zhang, Xuan Zhou 외

Recent progress in Multimodal Large Language Models (MLLMs) have significantly enhanced the ability of artificial intelligence systems to understand and generate multimodal content. However, these models often exhibit li…

BenchmarkingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models

2026-01-23 · Yuming Zhao, Peiyi Zhang, Oana Ignat arxiv

Memes are a pervasive form of online communication, yet their cultural specificity poses significant challenges for cross-cultural adaptation. We study cross-cultural meme transcreation, a multimodal generation task that…

multimodal generation

Can MLLMs Understand the Deep Implication Behind Chinese Images?

2024-10-17 · Chenhao Zhang, Xi Feng, Yuelin Bai, Xinrun Du 외

As the capabilities of Multimodal Large Language Models (MLLMs) continue to improve, the need for higher-order capability evaluation of MLLMs is increasing. However, there is a lack of work evaluating MLLM for higher-ord…

CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation

2024-07-01 · Yuxuan Wang, Yijun Liu, Fei Yu, Chen Huang 외

Despite the rapid development of Chinese vision-language models (VLMs), most existing Chinese vision-language (VL) datasets are constructed on Western-centric images from existing English VL datasets. The cultural bias i…

Image-text RetrievalQuestion AnsweringText RetrievalVisual Grounding+1