paper-with-me

홈 › Papers

Principia: Relational Physics Tests for Video Models

2026-09-03 · Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad hf

Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.

📄 PDF Abstract BibTeX arXiv:2609.04200

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reasoning over mathematical objects: on-policy reward modeling and test time aggregation

2026-03-19 · Pranjal Aggarwal, Marjan Ghazvininejad, Seungone Kim, Ilia Kulikov 외 arxiv

The ability to precisely derive mathematical objects is a core requirement for downstream STEM applications, including mathematics, physics, and chemistry, where reasoning must culminate in formally structured expression…

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

2025-05-29 · Xiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng 외

Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limit…

Self-Supervised LearningVideo GenerationVideo Understanding

Towards Robust Relational Causal Discovery

2019-12-05 · Sanghack Lee, Vasant Honavar

We consider the problem of learning causal relationships from relational data. Existing approaches rely on queries to a relational conditional independence (RCI) oracle to establish and orient causal relations in such a …

Causal Discovery

On Neural Architecture Inductive Biases for Relational Tasks

2022-06-09 · Giancarlo Kerg, Sarthak Mittal, David Rolnick, Yoshua Bengio 외

Current deep learning approaches have shown good in-distribution generalization performance, but struggle with out-of-distribution generalization. This is especially true in the case of tasks involving abstract relations…

Inductive BiasOut-of-Distribution Generalization

PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment

2026-03-14 · Zhexiao Xiong, Yizhi Song, Liu He, Wei Xiong 외 arxiv

Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing models often generate temporally incohe…

Synthetic Data GenerationPhysical IntuitionVideo Generation