paper-with-me

홈 › Papers

GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models

2023-11-15 · Serwan Jassim, Mario Holubar, Annika Richter, Cornelius Wolff, Xenia Ohmer, Elia Bruni

This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs). This evaluation is accomplished via a two-tier approach leveraging Unity simulations. The first level tests for language grounding by assessing a model's ability to relate simple textual descriptions with visual information. The second level evaluates the model's understanding of "Intuitive Physics" principles, such as object permanence and continuity. In addition to releasing the benchmark, we use it to evaluate several state-of-the-art multimodal LLMs. Our evaluation reveals significant shortcomings in the language grounding and intuitive physics capabilities of these models. Although they exhibit at least some grounding capabilities, particularly for colors and shapes, these capabilities depend heavily on the prompting strategy. At the same time, all models perform below or at the chance level of 50% in the Intuitive Physics tests, while human subjects are on average 80% correct. These identified limitations underline the importance of using benchmarks like GRASP to monitor the progress of future models in developing these competencies.

📄 PDF Abstract BibTeX arXiv:2311.09048

Code (0)

등록된 구현이 없습니다.

Tasks

Unity

Similar Papers 제목 키워드 기반

Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in Clutter

2023-11-09 · Georgios Tziafas, Yucheng Xu, Arushi Goel, Mohammadreza Kasaei 외

Robots operating in human-centric environments require the integration of visual grounding and grasping capabilities to effectively manipulate objects based on user instructions. This work focuses on the task of referrin…

ObjectVisual Grounding

Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views

2025-10-26 · Anna Deichler, Jonas Beskow arxiv

We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and stationary cameras, we recorded synchronize…

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking

2025-05-25 · Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari

Recent advances in large language models (LLMs) demonstrate their impressive reasoning capabilities. However, the reasoning confined to internal parametric space limits LLMs' access to real-time information and understan…

Mathematical ReasoningMulti-hop Question AnsweringQuestion Answeringtext-based games

RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation

2026-03-16 · Linfei Li, Lin Zhang, Ying Shen arxiv

Visual-language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existin…

Robotic Grasping

Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning

2025-09-09 · Houjian Yu, Zheming Zhou, Min Sun, Omid Ghasemalizadeh 외 arxiv

Enabling robots to grasp objects specified through natural language is essential for effective human-robot interaction, yet it remains a significant challenge. Existing approaches often struggle with open-form language e…

Spatial ReasoningRobotic Grasping