paper-with-me

Papers

VideoPhy: Evaluating Physical Commonsense for Video Generation

2024-06-05 · Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, Aditya Grover

Recent advances in internet-scale video data pretraining have led to the development of text-to-video generative models that can create high-quality videos across a broad range of visual concepts, synthesize realistic motions and render complex objects. Hence, these generative models have the potential to become general-purpose simulators of the physical world. However, it is unclear how far we are from this goal with the existing text-to-video generative models. To this end, we present VideoPhy, a benchmark designed to assess whether the generated videos follow physical commonsense for real-world activities (e.g. marbles will roll down when placed on a slanted surface). Specifically, we curate diverse prompts that involve interactions between various material types in the physical world (e.g., solid-solid, solid-fluid, fluid-fluid). We then generate videos conditioned on these captions from diverse state-of-the-art text-to-video generative models, including open models (e.g., CogVideoX) and closed models (e.g., Lumiere, Dream Machine). Our human evaluation reveals that the existing models severely lack the ability to generate videos adhering to the given text prompts, while also lack physical commonsense. Specifically, the best performing model, CogVideoX-5B, generates videos that adhere to the caption and physical laws for 39.6% of the instances. VideoPhy thus highlights that the video generative models are far from accurately simulating the physical world. Finally, we propose an auto-evaluator, VideoCon-Physics, to assess the performance reliably for the newly released models.

📄 PDF Abstract BibTeX arXiv:2406.03520

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

2025-03-09 · Hritik Bansal, Clark Peng, Yonatan Bitton, Roman Goldenberg 외

Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However, their adherence to physical commonsense …

Video Generation

NEWTON: Agentic Planning for Physically Grounded Video Generation

2026-05-18 · Yuxiang Feng, Juncheng Wang, Chao Xu, Yijie Qian 외 arxiv

Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves only 32.6% joint accuracy. We identify a specification bottleneck: tex…

Video Generation

PhysVid: Physics Aware Local Conditioning for Generative Video Models

2026-03-27 · Saurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram Zonooz arxiv

Generative video models achieve high visual fidelity but often violate basic physical principles, limiting reliability in real-world settings. Prior attempts to inject physics rely on conditioning: frame-level signals ar…

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation

2026-05-27 · Mariam Hassan, Kaouther Messaoud, Wuyang Li, Alexandre Alahi arxiv

Modern video generative models produce visually impressive results, yet frequently violate basic physical principles. We propose Proprio, a training-free framework that enables a frozen video generator to assess and impr…

Video Generation

Chain of Event-Centric Causal Thought for Physically Plausible Video Generation

2026-03-10 · Zixuan Wang, Yixin Hu, Haolan Wang, Feng Chen 외 arxiv

Physically Plausible Video Generation (PPVG) has emerged as a promising avenue for modeling real-world physical phenomena. PPVG requires an understanding of commonsense knowledge, which remains a challenge for video diff…

Video Generation