paper-with-me

Papers

Do generative video models understand physical principles?

2025-01-14 · Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, Robert Geirhos

AI video generation is undergoing a revolution, with quality and realism advancing rapidly. These advances have led to a passionate scientific debate: Do video models learn "world models" that discover laws of physics -- or, alternatively, are they merely sophisticated pixel predictors that achieve visual realism without understanding the physical principles of reality? We address this question by developing Physics-IQ, a comprehensive benchmark dataset that can only be solved by acquiring a deep understanding of various physical principles, like fluid dynamics, optics, solid mechanics, magnetism and thermodynamics. We find that across a range of current models (Sora, Runway, Pika, Lumiere, Stable Video Diffusion, and VideoPoet), physical understanding is severely limited, and unrelated to visual realism. At the same time, some test cases can already be successfully solved. This indicates that acquiring certain physical principles from observation alone may be possible, but significant challenges remain. While we expect rapid advances ahead, our work demonstrates that visual realism does not imply physical understanding. Our project page is at https://physics-iq.github.io; code at https://github.com/google-deepmind/physics-IQ-benchmark.

📄 PDF Abstract BibTeX arXiv:2501.09038

Code (1)

google-deepmind/physics-IQ-benchmark 공식 구현

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

GenMatter: Perceiving Physical Objects with Generative Matter Models

2026-04-24 · Eric Li, Arijit Dasgupta, Yoni Friedman, Mathieu Huot 외 arxiv

Human visual perception offers valuable insights for understanding computational principles of motion-based scene interpretation. Humans robustly detect and segment moving entities that constitute independently moveable …

Object SegmentationScene Understanding

PhysVid: Physics Aware Local Conditioning for Generative Video Models

2026-03-27 · Saurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram Zonooz arxiv

Generative video models achieve high visual fidelity but often violate basic physical principles, limiting reliability in real-world settings. Prior attempts to inject physics rely on conditioning: frame-level signals ar…

Physical Informed Driving World Model

2024-12-11 · Zhuoran Yang, Xi Guo, Chenjing Ding, Chiyu Wang 외

Autonomous driving requires robust perception models trained on high-quality, large-scale multi-view driving videos for tasks like 3D object detection, segmentation and trajectory prediction. While world models provide a…

3D Object DetectionAutonomous Drivingmodelobject-detection+3

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

2026-07-11 · Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu arxiv

Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object inte…

Physical Commonsense ReasoningVisual Question Answering

T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

2025-05-01 · Xuyang Guo, Jiayan Huo, Zhenmei Shi, Zhao Song 외

Text-to-video generative models have made significant strides in recent years, producing high-quality videos that excel in both aesthetic appeal and accurate instruction following, and have become central to digital art …

counterfactualInstruction FollowingText-to-Video GenerationVideo Generation