paper-with-me

홈 › Papers

PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models

2026-01-22 · Chak-Wing Mak, Guanyu Zhu, Boyi Zhang, Hongji Li, Xiaowei Chi, Kevin Zhang, Yichen Wu, Yangfan He, Chun-Kai Fan, Wentao Lu, Kuangzhi Ge, Xinyu Fang, Hongyang He, Kuan Lu, Tianxiang Xu, Li Zhang, Yongxin Ni, Youhua Li, Shanghang Zhang arxiv

Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp of the underlying physics remains underexplored. Existing benchmarks attempting to measure this matter rely on synthetic, Visual Question Answer templates or focus on perceptual video quality that is tangential to measuring how well the video abides by physical laws. To address this fragmentation, we introduce PhysicsMind, a unified benchmark with both real and simulation environments that evaluates law-consistent reasoning and generation over three canonical principles: Center of Mass, Lever Equilibrium, and Newton's First Law. PhysicsMind comprises two main tasks: i) VQA tasks, testing whether models can reason and determine physical quantities and values from images or short videos, and ii) Video Generation(VG) tasks, evaluating if predicted motion trajectories obey the same center-of-mass, torque, and inertial constraints as the ground truth. A broad range of recent models and video generation models is evaluated on PhysicsMind and found to rely on appearance heuristics while often violating basic mechanics. These gaps indicate that current scaling and training are still insufficient for robust physical understanding, underscoring PhysicsMind as a focused testbed for physics-aware multimodal models. Our data will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2601.16007

Code (0)

등록된 구현이 없습니다.

Tasks

Visual ReasoningVideo Generation

Similar Papers 제목 키워드 기반

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

2026-07-17 · Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang 외 hf

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without…

Video Generation

Which priors matter? Benchmarking models for learning latent dynamics

2021-11-09 · Aleksandar Botev, Andrew Jaegle, Peter Wirnsberger, Daniel Hennes 외

Learning dynamics is at the heart of many important applications of machine learning (ML), such as robotics and autonomous driving. In these settings, ML algorithms typically need to reason about a physical system using …

Autonomous DrivingBenchmarking

MECHBench: A Set of Black-Box Optimization Benchmarks originated from Structural Mechanics

2025-11-13 · Iván Olarte Rodríguez, Maria Laura Santoni, Fabian Duddeck, Carola Doerr 외 arxiv

Benchmarking is essential for developing and evaluating black-box optimization algorithms, providing a structured means to analyze their search behavior. Its effectiveness relies on carefully selected problem sets used f…

PhyGround: Benchmarking Physical Reasoning in Generative World Models

2026-05-11 · Juyi Lin, Arash Akbari, Yumei He, Lin Zhao 외 arxiv

Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, evaluating whether generated videos actual…

Video Generation

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs

2025-12-23 · Saeed Mohammadzadeh, Erfan Hamdi, Joel Shor, Emma Lejeune arxiv

As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to generate scientifically valid physical models has become a critical gap. Computati…