paper-with-me

Papers

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?

2026-05-11 · Maximilian Triebel, Marco Menner, Dominik Helfenstein arxiv

Vision-Language(-Action) Models (VLMs) are increasingly applied to interactive environments, yet existing benchmarks often overlook the complex physical reasoning required for point-and-click puzzle games. This paper introduces Vision-Language Against The Incredible Machine (VLATIM), a benchmark designed to evaluate human-like logical problem-solving capabilities within the classic physics puzzle game The Incredible Machine 2 (TIM). Unlike existing benchmarks, VLATIM specifically targets the critical gap between high-level logical reasoning and continuous action spaces requiring precise mouse interactions. This benchmark is structured into five progressive parts, assessing capabilities that range from basic visual grounding and domain understanding to multi-step manipulation and full puzzle solving. Our results reveal a significant disparity between reasoning and execution. While large proprietary models demonstrate superior planning abilities, they struggle with precise visual grounding. Consequently, they do not yet show human-like problem-solving capabilities.

📄 PDF Abstract BibTeX arXiv:2605.11223

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningVisual Grounding

Similar Papers 제목 키워드 기반

Dual Thinking and Logical Processing -- Are Multi-modal Large Language Models Closing the Gap with Human Vision ?

2024-06-11 · Kailas Dayanandan, Nikhil Kumar, Anand Sinha, Brejesh lall

The dual thinking framework considers fast, intuitive, and slower logical processing. The perception of dual thinking in vision requires images where inferences from intuitive and logical processing differ, and the latte…

Autonomous DrivingDeep LearningInstance SegmentationLogical Reasoning+2

Learning to Learn Semantic Parsers from Natural Language Supervision

2019-02-22 · EMNLP 2018 10 · Igor Labutov, Bishan Yang, Tom Mitchell

As humans, we often rely on language to learn language. For example, when corrected in a conversation, we may learn from that correction, over time improving our language fluency. Inspired by this observation, we propose…

Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts

2025-05-19 · Zekun Wang, Sashank Varma

With the rapid improvement of machine learning (ML) models, cognitive scientists are increasingly asking about their alignment with how humans think. Here, we ask this question for computer vision models and human sensit…

Odd One OutSensitivity

How Well Do Deep Learning Models Capture Human Concepts? The Case of the Typicality Effect

2024-05-25 · Siddhartha K. Vemuri, Raj Sanjay Shah, Sashank Varma

How well do representations learned by ML models align with those of humans? Here, we consider concept representations learned by deep learning models and evaluate whether they show a fundamental behavioral signature of …

Language ModelingLanguage Modelling

Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding

2026-05-18 · Shravan Murlidaran, Ziqi Wen, Sana Shehabi, Miguel P. Eckstein arxiv

When humans view scenes without a specific task (free-viewing), they initially direct their eye movements toward the scene center and then fixate on people, text, objects being gazed at or grasped, and semantically meani…

Scene Understanding