paper-with-me

홈 › Papers

Comparing Visual Reasoning in Humans and AI

2021-04-29 · Shravan Murlidaran, William Yang Wang, Miguel P. Eckstein

Recent advances in natural language processing and computer vision have led to AI models that interpret simple scenes at human levels. Yet, we do not have a complete understanding of how humans and AI models differ in their interpretation of more complex scenes. We created a dataset of complex scenes that contained human behaviors and social interactions. AI and humans had to describe the scenes with a sentence. We used a quantitative metric of similarity between scene descriptions of the AI/human and ground truth of five other human descriptions of each scene. Results show that the machine/human agreement scene descriptions are much lower than human/human agreement for our complex scenes. Using an experimental manipulation that occludes different spatial regions of the scenes, we assessed how machines and humans vary in utilizing regions of images to understand the scenes. Together, our results are a first step toward understanding how machines fall short of human visual reasoning with complex scenes depicting human behaviors.

📄 PDF Abstract BibTeX arXiv:2104.14102

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceVisual Reasoning

Similar Papers 제목 키워드 기반

Five Points to Check when Comparing Visual Perception in Humans and Machines

2020-04-20 · Christina M. Funke, Judy Borowski, Karolina Stosio, Wieland Brendel 외

With the rise of machines to human-level performance in complex recognition tasks, a growing amount of work is directed towards comparing information processing in humans and machines. These studies are an exciting chanc…

Decision MakingObject RecognitionVisual Reasoning

VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs

2025-07-04 · Shmuel Berman, Jia Deng arxiv

Vision-Language Models (VLMs) excel at complex visual tasks such as VQA and chart understanding, yet recent work suggests they struggle with simple perceptual tests. We present an evaluation of vision-language models' ca…

Visual Reasoning

On the Cognition of Visual Question Answering Models and Human Intelligence: A Comparative Study

2023-10-04 · Liben Chen, Long Chen, Tian Ellison-Chen, Zhuoyuan Xu

Visual Question Answering (VQA) is a challenging task that requires cross-modal understanding and reasoning of visual image and natural language question. To inspect the association of VQA models to human cognition, we d…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs

2026-05-18 · Tara Azin, Yongan Yu, Raj Singh, Olessia Jouravlev arxiv

Presupposition projection in conditionals is central to theories of meaning and pragmatics, yet it remains largely unevaluated in large language models. We address this gap through a parallel behavioral study comparing h…

iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs

2025-02-05 · Julius Mayer, Mohamad Ballout, Serwan Jassim, Farbod Nosrat Nezami 외

Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multi-modal benchmark designed to evaluate the spa…

Spatial Reasoning