paper-with-me

Papers

IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models

2024-03-23 · HAZ Sameen Shahgir, Khondker Salman Sayeed, Abhik Bhattacharjee, Wasi Uddin Ahmad, Yue Dong, Rifat Shahriyar

The advent of Vision Language Models (VLM) has allowed researchers to investigate the visual understanding of a neural network using natural language. Beyond object classification and detection, VLMs are capable of visual comprehension and common-sense reasoning. This naturally led to the question: How do VLMs respond when the image itself is inherently unreasonable? To this end, we present IllusionVQA: a diverse dataset of challenging optical illusions and hard-to-interpret scenes to test the capability of VLMs in two distinct multiple-choice VQA tasks - comprehension and soft localization. GPT4V, the best performing VLM, achieves 62.99% accuracy (4-shot) on the comprehension task and 49.7% on the localization task (4-shot and Chain-of-Thought). Human evaluation reveals that humans achieve 91.03% and 100% accuracy in comprehension and localization. We discover that In-Context Learning (ICL) and Chain-of-Thought reasoning substantially degrade the performance of Gemini-Pro in the localization task. Tangentially, we discover a potential weakness in the ICL capabilities of VLMs: they fail to locate optical illusions even when the correct answer is in the context window as a few-shot example.

📄 PDF Abstract BibTeX arXiv:2403.15952

Code (1)

csebuetnlp/illusionvqa 공식 구현 pytorch

Tasks

Common Sense ReasoningIn-Context LearningMultiple-choiceObject LocalizationVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions

2025-07-30 · Yiting Qu, Ziqing Yang, Yihan Ma, Michael Backes 외 arxiv

Recent advances in text-to-image diffusion models have enabled the creation of a new form of digital art: optical illusions--visual tricks that create different perceptions of reality. However, adversaries may misuse suc…

Informing Computer Vision with Optical Illusions

2019-02-08 · Nasim Nematzadeh, David M. W. Powers, Trent Lewis

Illusions are fascinating and immediately catch people's attention and interest, but they are also valuable in terms of giving us insights into human cognition and perception. A good theory of human perception should be …

Do vision models perceive illusory motion in static images like humans?

2026-04-10 · Isabella Elaine Rosario, Fan L. Cheng, Zitang Sun, Nikolaus Kriegeskorte arxiv

Understanding human motion processing is essential for building reliable, human-centered computer vision systems. Although deep neural networks (DNNs) achieve strong performance in optical flow estimation, they remain le…

Optical Flow Estimation

Optical Illusions Images Dataset

2018-09-30 · Robert Max Williams, Roman V. Yampolskiy

Human vision is capable of performing many tasks not optimized for in its long evolution. Reading text and identifying artificial objects such as road signs are both tasks that mammalian brains never encountered in the w…

Seeing the Evidence, Missing the Answer: Tool-Guided Vision-Language Models on Visual Illusions

2026-03-31 · Xuesong Wang, Harry Wang arxiv

Vision-language models (VLMs) exhibit a systematic bias when confronted with classic optical illusions: they overwhelmingly predict the illusion as "real" regardless of whether the image has been counterfactually modifie…

Image ManipulationSpatial ReasoningImage Compression