paper-with-me

홈 › Papers

Benchmarks for Physical Reasoning AI

2023-12-17 · Andrew Melnik, Robin Schiewer, Moritz Lange, Andrei Muresanu, Mozhgan Saeidi, Animesh Garg, Helge Ritter

Physical reasoning is a crucial aspect in the development of general AI systems, given that human learning starts with interacting with the physical world before progressing to more complex concepts. Although researchers have studied and assessed the physical reasoning of AI approaches through various specific benchmarks, there is no comprehensive approach to evaluating and measuring progress. Therefore, we aim to offer an overview of existing benchmarks and their solution approaches and propose a unified perspective for measuring the physical reasoning capacity of AI systems. We select benchmarks that are designed to test algorithmic performance in physical reasoning tasks. While each of the selected benchmarks poses a unique challenge, their ensemble provides a comprehensive proving ground for an AI generalist agent with a measurable skill level for various physical reasoning concepts. This gives an advantage to such an ensemble of benchmarks over other holistic benchmarks that aim to simulate the real world by intertwining its complexity and many concepts. We group the presented set of physical reasoning benchmarks into subcategories so that more narrow generalist AI agents can be tested first on these groups.

📄 PDF Abstract BibTeX arXiv:2312.10728

Code (1)

ndrwmlnk/awesome-benchmarks-for-physical-reasoning-ai 공식 구현

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

2026-07-11 · Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu arxiv

Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object inte…

Physical Commonsense ReasoningVisual Question Answering

Hi-Phy: A Benchmark for Hierarchical Physical Reasoning

2021-06-17 · Cheng Xue, Vimukthini Pinto, Chathura Gamage, Peng Zhang 외

Reasoning about the behaviour of physical objects is a key capability of agents operating in physical worlds. Humans are very experienced in physical reasoning while it remains a major challenge for AI. To facilitate res…

Everyday Physics in Korean Contexts: A Culturally Grounded Physical Reasoning Benchmark

2025-09-22 · Jihae Jeong, DaeYeop Lee, DongGeon Lee, Hwanjo Yu arxiv

Existing physical commonsense reasoning benchmarks predominantly focus on Western contexts, overlooking cultural variations in physical problem-solving. To address this gap, we introduce EPiK (Everyday Physics in Korean …

Physical Commonsense Reasoning

ContPhy: Continuum Physical Concept Learning and Reasoning from Videos

2024-02-09 · Zhicheng Zheng, Xin Yan, Zhenfang Chen, Jingzhou Wang 외

We introduce the Continuum Physical Dataset (ContPhy), a novel benchmark for assessing machine physical commonsense. ContPhy complements existing physical reasoning benchmarks by encompassing the inference of diverse phy…

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning

2025-09-12 · Gautam Sreekumar, Vishnu Naresh Boddeti arxiv

Large multimodal models (LMMs) encode physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs to answer physical reasoning queries, such as the outcome of a potentia…

Visual Question Answering