paper-with-me

Papers

Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint

2025-05-29 · HeeKyung Lee, Jiaxin Ge, Tsung-Han Wu, Minwoo Kang, Trevor Darrell, David M. Chan

Rebus puzzles, visual riddles that encode language through imagery, spatial arrangement, and symbolic substitution, pose a unique challenge to current vision-language models (VLMs). Unlike traditional image captioning or question answering tasks, rebus solving requires multi-modal abstraction, symbolic reasoning, and a grasp of cultural, phonetic and linguistic puns. In this paper, we investigate the capacity of contemporary VLMs to interpret and solve rebus puzzles by constructing a hand-generated and annotated benchmark of diverse English-language rebus puzzles, ranging from simple pictographic substitutions to spatially-dependent cues ("head" over "heels"). We analyze how different VLMs perform, and our findings reveal that while VLMs exhibit some surprising capabilities in decoding simple visual clues, they struggle significantly with tasks requiring abstract reasoning, lateral thinking, and understanding visual metaphors.

📄 PDF Abstract BibTeX arXiv:2505.23759

Code (1)

kyunnilee/visual_puzzles 공식 구현

Tasks

Image CaptioningQuestion Answering

Similar Papers 제목 키워드 기반

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles

2025-08-02 · Yelim Ahn, Jaejin Lee arxiv

As large language models (LLMs) are increasingly deployed across diverse domains, ensuring their safety has become a critical concern. In response, studies on jailbreak attacks have been actively growing. Existing approa…

Prompt Engineering

Get Your Model Puzzled: Introducing Crossword-Solving as a New NLP Benchmark

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Solving crossword puzzles requires diverse reasoning capabilities, access to a vast amount of knowledge about language and the world, and the ability to satisfy the constraints imposed by the structure of the puzzle. In …

Natural Language UnderstandingOpen-Domain Question AnsweringQuestion AnsweringRetrieval

On Memorization of Large Language Models in Logical Reasoning

2024-10-30 · Chulin Xie, Yangsibo Huang, Chiyuan Zhang, Da Yu 외

Large language models (LLMs) achieve good performance on challenging reasoning benchmarks, yet could also make basic reasoning mistakes. This contrasting behavior is puzzling when it comes to understanding the mechanisms…

Logical ReasoningMemorization

Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles

2024-09-16 · Kulin Shah, Nishanth Dikkala, Xin Wang, Rina Panigrahy

Causal language modeling using the Transformer architecture has yielded remarkable capabilities in Large Language Models (LLMs) over the last few years. However, the extent to which fundamental search and reasoning capab…

Causal Language ModelingLanguage ModelingLanguage ModellingLogical Sequence

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

2026-06-30 · Mohammadamin Shafiei, Shuyue Stella Li, Yulia Tsvetkov arxiv

As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behaviors are genuine or superficial. We show that current fairness evaluat…