paper-with-me

홈 › Papers

Inference-time Physics Alignment of Video Generative Models with Latent World Models

2026-01-15 · Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez, Melissa Hall, Reyhane Askari-Hemmat, Xiaochuang Han, Nicolas Ballas, Michal Drozdzal, Adriana Romero-Soriano arxiv

State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we find that the shortfall in physics plausibility also stems from suboptimal inference strategies. We therefore introduce WMReward and treat improving physics plausibility of video generation as an inference-time alignment problem. In particular, we leverage the strong physics prior of a latent world model (here, VJEPA-2) as a reward to search and steer multiple candidate denoising trajectories, enabling scaling test-time compute for better generation performance. Empirically, our approach substantially improves physics plausibility across image-conditioned, multiframe-conditioned, and text-conditioned generation settings, with validation from human preference study. Notably, in the ICCV 2025 Perception Test PhysicsIQ Challenge, we achieve a final score of 62.64%, winning first place and outperforming the previous state of the art by 7.42%. Our work demonstrates the viability of using latent world models to improve physics plausibility of video generation, beyond this specific instantiation or parameterization.

📄 PDF Abstract BibTeX arXiv:2601.10553

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

PhysVid: Physics Aware Local Conditioning for Generative Video Models

2026-03-27 · Saurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram Zonooz arxiv

Generative video models achieve high visual fidelity but often violate basic physical principles, limiting reliability in real-world settings. Prior attempts to inject physics rely on conditioning: frame-level signals ar…

LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference

2025-10-13 · Jianhao Yuan, Fabio Pizzati, Francesco Pinto, Lars Kunze 외 arxiv

Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due …

Improving the Physics of Video Generation with VJEPA-2 Reward Signal

2025-10-22 · Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez 외 arxiv

This is a short technical report describing the winning entry of the PhysicsIQ Challenge, presented at the Perception Test Workshop at ICCV 2025. State-of-the-art video generative models exhibit severely limited physical…

Video Generation

PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction

2026-05-19 · Xin Zhang, Yabo Chen, Yijie Fang, Wanying Qu 외 arxiv

Recent generative video models achieve impressive visual quality but remain constrained by limited physical consistency and controllability. Existing video generation methods provide minimal physical control, and single-…

3D ReconstructionScene GenerationVideo Generation3D Generation

Beyond Rigid: Benchmarking Non-Rigid Video Editing

2026-01-26 · Bingzheng Qu, Xuefeng Bai, Kehai Chen, Min Zhang arxiv

As video generation models are increasingly expected to manipulate physical dynamics, there is a growing need to move evaluation beyond appearance fidelity and semantic alignment. Non-rigid video editing offers a uniquel…

Instruction FollowingVideo Generation