paper-with-me

홈 › Papers

Hi-Phy: A Benchmark for Hierarchical Physical Reasoning

2021-06-17 · Cheng Xue, Vimukthini Pinto, Chathura Gamage, Peng Zhang, Jochen Renz

Reasoning about the behaviour of physical objects is a key capability of agents operating in physical worlds. Humans are very experienced in physical reasoning while it remains a major challenge for AI. To facilitate research addressing this problem, several benchmarks have been proposed recently. However, these benchmarks do not enable us to measure an agent's granular physical reasoning capabilities when solving a complex reasoning task. In this paper, we propose a new benchmark for physical reasoning that allows us to test individual physical reasoning capabilities. Inspired by how humans acquire these capabilities, we propose a general hierarchy of physical reasoning capabilities with increasing complexity. Our benchmark tests capabilities according to this hierarchy through generated physical reasoning tasks in the video game Angry Birds. This benchmark enables us to conduct a comprehensive agent evaluation by measuring the agent's granular physical reasoning capabilities. We conduct an evaluation with human players, learning agents, and heuristic agents and determine their capabilities. Our evaluation shows that learning agents, with good local generalization ability, still struggle to learn the underlying physical reasoning capabilities and perform worse than current state-of-the-art heuristic agents and humans. We believe that this benchmark will encourage researchers to develop intelligent agents with advanced, human-like physical reasoning capabilities. URL: https://github.com/Cheng-Xue/Hi-Phy

📄 PDF Abstract BibTeX arXiv:2106.09692

Code (1)

Cheng-Xue/Hi-Phy 공식 구현 pytorch

Similar Papers 제목 키워드 기반

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

2026-07-11 · Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu arxiv

Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object inte…

Physical Commonsense ReasoningVisual Question Answering

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

2025-03-18 · Nvidia, :, Alisson Azzolini, Junjie Bai 외

Physical AI systems need to perceive, understand, and perform complex actions in the physical world. In this paper, we present the Cosmos-Reason1 models that can understand the physical world and generate appropriate emb…

3D Face AnimationCommon Sense ReasoningReinforcement Learning (RL)

Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models

2025-11-23 · Tianyang Han, Junhao Su, Junjie Hu, Peizhen Yang 외 arxiv

Text-to-image (T2I) models today are capable of producing photorealistic, instruction-following images, yet they still frequently fail on prompts that require implicit world knowledge. Existing evaluation protocols eithe…

SPHERE: A Hierarchical Evaluation on Spatial Perception and Reasoning for Vision-Language Models

2024-12-17 · Wenyu Zhang, Wei En Ng, Lixin Ma, Yuwen Wang 외

Current vision-language models may incorporate single-dimensional spatial cues, such as depth, object boundary, and basic spatial directions (e.g. left, right, front, back), yet often lack the multi-dimensional spatial r…

Logical ReasoningSpatial Reasoning

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

2026-06-25 · Shengbin Guo, Shaokang He, Chaoyue Meng, Shengpeng Xiao 외 arxiv

While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing benchmarks lack a comprehensive evaluation of physics-based reasoning, a critical capability for handl…

Video GenerationImage Editing