paper-with-me

Papers

The Illusion-Illusion: Vision Language Models See Illusions Where There are None

2024-12-07 · Tomer Ullman

Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something "really is" and how something "appears to be", and this gap helps us understand the mental processing that lead to how something appears to be. Illusions are also useful for investigating artificial systems, and much research has examined whether computational models of perceptions fall prey to the same illusions as people. Here, I invert the standard use of perceptual illusions to examine basic processing errors in current vision language models. I present these models with illusory-illusions, neighbors of common illusions that should not elicit processing errors. These include such things as perfectly reasonable ducks, crooked lines that truly are crooked, circles that seem to have different sizes because they are, in fact, of different sizes, and so on. I show that many current vision language systems mistakenly see these illusion-illusions as illusions. I suggest that such failures are part of broader failures already discussed in the literature.

📄 PDF Abstract BibTeX arXiv:2412.18613

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticPhilosophy

Similar Papers 제목 키워드 기반

Evaluating Model Perception of Color Illusions in Photorealistic Scenes

2024-12-09 · CVPR 2025 1 · Lingjun Mao, Zineng Tang, Alane Suhr

We study the perception of color illusions by vision-language models. Color illusion, where a person's visual system perceives color differently from actual color, is well-studied in human vision. However, it remains und…

Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?

2023-10-31 · Yichi Zhang, Jiayi Pan, Yuchen Zhou, Rui Pan 외

Vision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world. However, known as visual illusions, human's perception of reality isn't always faithful to th…

IllusionBench: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models

2025-01-01 · Yiming Zhang, ZiCheng Zhang, Xinyi Wei, Xiaohong Liu 외

Current Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have bee…

HallucinationMultiple-choice

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

2026-02-02 · Wenjin Hou, Wei Liu, Han Hu, Xiaoxiao Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. However, these evaluations typically rely on s…

Visual Reasoning

Synthesizing Visual Illusions Using Generative Adversarial Networks

2019-11-21 · Alexander Gomez-Villa, Adrian Martín, Javier Vazquez-Corral, Jesús Malo 외

Visual illusions are a very useful tool for vision scientists, because they allow them to better probe the limits, thresholds and errors of the visual system. In this work we introduce the first ever framework to generat…

Generative Adversarial Network