paper-with-me

Papers

VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

2025-03-09 · Hritik Bansal, Clark Peng, Yonatan Bitton, Roman Goldenberg, Aditya Grover, Kai-Wei Chang

Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However, their adherence to physical commonsense across real-world actions remains unclear (e.g., playing tennis, backflip). Existing benchmarks suffer from limitations such as limited size, lack of human evaluation, sim-to-real gaps, and absence of fine-grained physical rule analysis. To address this, we introduce VideoPhy-2, an action-centric dataset for evaluating physical commonsense in generated videos. We curate 200 diverse actions and detailed prompts for video synthesis from modern generative models. We perform human evaluation that assesses semantic adherence, physical commonsense, and grounding of physical rules in the generated videos. Our findings reveal major shortcomings, with even the best model achieving only 22% joint performance (i.e., high semantic and physical commonsense adherence) on the hard subset of VideoPhy-2. We find that the models particularly struggle with conservation laws like mass and momentum. Finally, we also train VideoPhy-AutoEval, an automatic evaluator for fast, reliable assessment on our dataset. Overall, VideoPhy-2 serves as a rigorous benchmark, exposing critical gaps in video generative models and guiding future research in physically-grounded video generation. The data and code is available at https://videophy2.github.io/.

📄 PDF Abstract BibTeX arXiv:2503.06800

Code (1)

Hritikbansal/videophy/tree/main/VIDEOPHY2 pytorch

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VideoPhy: Evaluating Physical Commonsense for Video Generation

2024-06-05 · Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong 외

Recent advances in internet-scale video data pretraining have led to the development of text-to-video generative models that can create high-quality videos across a broad range of visual concepts, synthesize realistic mo…

Video Generation

MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models

2025-10-21 · Aritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek, Amirhossein Habibian 외 arxiv

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex …

PhysVid: Physics Aware Local Conditioning for Generative Video Models

2026-03-27 · Saurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram Zonooz arxiv

Generative video models achieve high visual fidelity but often violate basic physical principles, limiting reliability in real-world settings. Prior attempts to inject physics rely on conditioning: frame-level signals ar…

NEWTON: Agentic Planning for Physically Grounded Video Generation

2026-05-18 · Yuxiang Feng, Juncheng Wang, Chao Xu, Yijie Qian 외 arxiv

Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves only 32.6% joint accuracy. We identify a specification bottleneck: tex…

Video Generation

Chain of Event-Centric Causal Thought for Physically Plausible Video Generation

2026-03-10 · Zixuan Wang, Yixin Hu, Haolan Wang, Feng Chen 외 arxiv

Physically Plausible Video Generation (PPVG) has emerged as a promising avenue for modeling real-world physical phenomena. PPVG requires an understanding of commonsense knowledge, which remains a challenge for video diff…

Video Generation