paper-with-me

홈 › Papers

Reasoning Riddles: How Explainability Reveals Cognitive Limits in Vision-Language Models

2025-10-03 · Prahitha Movva arxiv

Vision-Language Models (VLMs) excel at many multimodal tasks, yet their cognitive processes remain opaque on complex lateral thinking challenges like rebus puzzles. While recent work has demonstrated these models struggle significantly with rebus puzzle solving, the underlying reasoning processes and failure patterns remain largely unexplored. We address this gap through a comprehensive explainability analysis that moves beyond performance metrics to understand how VLMs approach these complex lateral thinking challenges. Our study contributes a systematically annotated dataset of 221 rebus puzzles across six cognitive categories, paired with an evaluation framework that separates reasoning quality from answer correctness. We investigate three prompting strategies designed to elicit different types of explanatory processes and reveal critical insights into VLM cognitive processes. Our findings demonstrate that reasoning quality varies dramatically across puzzle categories, with models showing systematic strengths in visual composition while exhibiting fundamental limitations in absence interpretation and cultural symbolism. We also discover that prompting strategy substantially influences both cognitive approach and problem-solving effectiveness, establishing explainability as an integral component of model performance rather than a post-hoc consideration.

📄 PDF Abstract BibTeX arXiv:2510.02780

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

2025-09-18 · Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Jiyuan Fang 외 arxiv

Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark p…

The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans

2026-06-25 · Bella Fascendini, Kathryn McGregor, Max D. Gupta, Thomas L. Griffiths arxiv

Humans flexibly adapt their reasoning strategies to the requirements of a given problem. Large language models (LLMs) have performed well on many cognitive tasks, however, it is unclear whether this accuracy is a result …

Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models

2024-07-28 · Nitzan Bitton-Guetta, Aviv Slobodkin, Aviya Maimon, Eliya Habba 외

Imagine observing someone scratching their arm; to understand why, additional context would be necessary. However, spotting a mosquito nearby would immediately offer a likely explanation for the person's discomfort, ther…

World Knowledge

RiddleSense: Reasoning about Riddle Questions Featuring Linguistic Creativity and Commonsense Knowledge

2021-01-02 · Findings (ACL) 2021 8 · Bill Yuchen Lin, Ziyi Wu, Yichi Yang, Dong-Ho Lee 외

Question: I have five fingers but I am not alive. What am I? Answer: a glove. Answering such a riddle-style question is a challenging cognitive process, in that it requires complex commonsense reasoning abilities, an und…

counterfactualCounterfactual ReasoningMultiple-choiceNatural Language Understanding+1

ICL Optimized Fragility

2025-09-30 · Serena Gomez Wannaz arxiv

ICL guides are known to improve task-specific performance, but their impact on cross-domain cognitive abilities remains unexplored. This study examines how ICL guides affect reasoning across different knowledge domains u…

Mathematical ReasoningGeneral Knowledge