paper-with-me

Papers

I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes

2026-03-24 · Shijia Zhou, Saif M. Mohammad, Barbara Plank, Diego Frassinelli arxiv

Internet memes represent a popular form of multimodal online communication and often use figurative elements to convey layered meaning through the combination of text and images. However, it remains largely unclear how multimodal large language models (MLLMs) combine and interpret visual and textual information to identify figurative meaning in memes. To address this gap, we evaluate eight state-of-the-art generative MLLMs across three datasets on their ability to detect and explain six types of figurative meaning. In addition, we conduct a human evaluation of the explanations generated by these MLLMs, assessing whether the provided reasoning supports the predicted label and whether it remains faithful to the original meme content. Our findings indicate that all models exhibit a strong bias to associate a meme with figurative meaning, even when no such meaning is present. Qualitative analysis further shows that correct predictions are not always accompanied by faithful explanations.

📄 PDF Abstract BibTeX arXiv:2603.23229

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework

2025-11-28 · Kelaiti Xiao, Liang Yang, Dongyu Zhang, Paerhati Tulajiang 외 arxiv

We study idiom-based visual puns--images that align an idiom's literal and figurative meanings--and present an iterative framework that coordinates a large language model (LLM), a text-to-image model (T2IM), and a multim…

IRFL: Image Recognition of Figurative Language

2023-03-27 · Ron Yosef, Yonatan Bitton, Dafna Shahaf

Figures of speech such as metaphors, similes, and idioms are integral parts of human communication. They are ubiquitous in many forms of discourse, allowing people to convey complex, abstract ideas and evoke emotion. As …

ClassificationVisual Reasoning

Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding

2026-01-23 · Seyyed Saeid Cheshmi, Hahnemann Ortiz, James Mooney, Dongyeop Kang arxiv

Vision-language models (VLMs) have demonstrated strong reasoning abilities in literal multimodal tasks such as visual mathematics and science question answering. However, figurative language, such as sarcasm, humor, and …

Science Question AnsweringMultimodal Reasoning

FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language

2026-01-27 · Faezeh Hosseini, Mohammadali Yousefzadeh, Yadollah Yaghoobzadeh arxiv

Figurative language, particularly fixed figurative expressions (FFEs) such as idioms and proverbs, poses persistent challenges for large language models (LLMs). Unlike literal phrases, FFEs are culturally grounded, large…

Exploring Concreteness Through a Figurative Lens

2026-04-20 · Saptarshi Ghosh, Tianyu Jiang arxiv

Static concreteness ratings are widely used in NLP, yet a word's concreteness can shift with context, especially in figurative language such as metaphor, where common concrete nouns can take abstract interpretations. Whi…