paper-with-me

Papers

PhyRPR: Training-Free Physics-Constrained Video Generation

2026-01-14 · Yibo Zhao, Hengjia Li, Xiaofei He, Boxi Wu arxiv

Recent diffusion-based video generation models can synthesize visually plausible videos, yet they often struggle to satisfy physical constraints. A key reason is that most existing approaches remain single-stage: they entangle high-level physical understanding with low-level visual synthesis, making it hard to generate content that require explicit physical reasoning. To address this limitation, we propose a training-free three-stage pipeline,\textit{PhyRPR}:\textit{Phy\uline{R}eason}--\textit{Phy\uline{P}lan}--\textit{Phy\uline{R}efine}, which decouples physical understanding from visual synthesis. Specifically, \textit{PhyReason} uses a large multimodal model for physical state reasoning and an image generator for keyframe synthesis; \textit{PhyPlan} deterministically synthesizes a controllable coarse motion scaffold; and \textit{PhyRefine} injects this scaffold into diffusion sampling via a latent fusion strategy to refine appearance while preserving the planned dynamics. This staged design enables explicit physical control during generation. Extensive experiments under physics constraints show that our method consistently improves physical plausibility and motion controllability.

📄 PDF Abstract BibTeX arXiv:2601.09255

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation

2025-04-21 · Weijie He, Mushui Liu, Yunlong Yu, Zhao Wang 외

Compositional text-to-video generation, which requires synthesizing dynamic scenes with multiple interacting entities and precise spatial-temporal relationships, remains a critical challenge for diffusion-based models. E…

AttributeDenoisingText-to-Video GenerationVideo Alignment+1

FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity

2025-06-09 · CVPR 2025 1 · Jinxi Li, Ziyang Song, Siyuan Zhou, Bo Yang

In this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural ne…

Motion Segmentation

A physics-encoded Fourier neural operator approach for surrogate modeling of divergence-free stress fields in solids

2024-08-27 · Mohammad S. Khorrami, Pawan Goyal, Jaber R. Mianroodi, Bob Svendsen 외

The purpose of the current work is the development of a so-called physics-encoded Fourier neural operator (PeFNO) for surrogate modeling of the quasi-static equilibrium stress field in solids. Rather than accounting for …

PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction

2026-05-19 · Xin Zhang, Yabo Chen, Yijie Fang, Wanying Qu 외 arxiv

Recent generative video models achieve impressive visual quality but remain constrained by limited physical consistency and controllability. Existing video generation methods provide minimal physical control, and single-…

3D ReconstructionScene GenerationVideo Generation3D Generation

Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement

2025-11-25 · Yang Liu, Xilin Zhao, Peisong Wen, Siran Dai 외 arxiv

Recent progress in video generation has led to impressive visual quality, yet current models still struggle to produce results that align with real-world physical principles. To this end, we propose an iterative self-ref…

Video Generation