paper-with-me

Papers

Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!

2024-10-01 · Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee, Youngjae Yu

Humans possess multimodal literacy, allowing them to actively integrate information from various modalities to form reasoning. Faced with challenges like lexical ambiguity in text, we supplement this with other modalities, such as thumbnail images or textbook illustrations. Is it possible for machines to achieve a similar multimodal understanding capability? In response, we present Understanding Pun with Image Explanations (UNPIE), a novel benchmark designed to assess the impact of multimodal inputs in resolving lexical ambiguities. Puns serve as the ideal subject for this evaluation due to their intrinsic ambiguity. Our dataset includes 1,000 puns, each accompanied by an image that explains both meanings. We pose three multimodal challenges with the annotations to assess different aspects of multimodal literacy; Pun Grounding, Disambiguation, and Reconstruction. The results indicate that various Socratic Models and Visual-Language Models improve over the text-only models when given visual context, particularly as the complexity of the tasks increases.

📄 PDF Abstract BibTeX arXiv:2410.01023

Code (1)

jiwanchung/visualpun_unpie 공식 구현

Similar Papers 제목 키워드 기반

LaViSA: A Language and Vision Structural Ambiguity Benchmark

2026-06-17 · Lee Sangmyeong, Shun Inadumi, Koichiro Yoshino arxiv

Structural ambiguity arises when a single sentence admits multiple valid interpretations due to its syntactic structure, posing a fundamental challenge for language understanding. Visual scenes serve as useful cues for r…

Visual Word Sense Disambiguation with CLIP through Dual-Channel Text Prompting and Image Augmentations

2026-02-06 · Shamik Bhattacharya, Daniel Perkins, Yaren Dogan, Vineeth Konjeti 외 arxiv

Ambiguity poses persistent challenges in natural language understanding for large language models (LLMs). To better understand how lexical ambiguity can be resolved through the visual domain, we develop an interpretable …

Natural Language UnderstandingWord Sense DisambiguationImage Augmentation

Point What You Mean: Visually Grounded Instruction Policy

2025-12-22 · Hang Yu, Juntu Zhao, Yufeng Liu, Kaiyu Li 외 arxiv

Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especially in cluttered or out-of-distribution (…

Visual Grounding

Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

2026-06-30 · Ta Duc Huy, Trang Nguyen, Townim Chowdhury, Ankit Yadav 외 arxiv

Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such as Semantic Entropy (SE), rely on output diversity. Yet our analysis …

DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech

2025-06-09 · Haotian Guo, Jing Han, Yongfeng Tu, Shihao Gao 외

Despite extensive research on textual and visual disambiguation, disambiguation through speech (DTS) remains underexplored. This is largely due to the lack of high-quality datasets that pair spoken sentences with richly …