paper-with-me

홈 › Papers

Think Before You Diffuse: LLMs-Guided Physics-Aware Video Generation

2025-05-27 · Ke Zhang, Cihan Xiao, Yiqun Mei, Jiacong Xu, Vishal M. Patel

Recent video diffusion models have demonstrated their great capability in generating visually-pleasing results, while synthesizing the correct physical effects in generated videos remains challenging. The complexity of real-world motions, interactions, and dynamics introduce great difficulties when learning physics from data. In this work, we propose DiffPhy, a generic framework that enables physically-correct and photo-realistic video generation by fine-tuning a pre-trained video diffusion model. Our method leverages large language models (LLMs) to explicitly reason a comprehensive physical context from the text prompt and use it to guide the generation. To incorporate physical context into the diffusion model, we leverage a Multimodal large language model (MLLM) as a supervisory signal and introduce a set of novel training objectives that jointly enforce physical correctness and semantic consistency with the input text. We also establish a high-quality physical video dataset containing diverse phyiscal actions and events to facilitate effective finetuning. Extensive experiments on public benchmarks demonstrate that DiffPhy is able to produce state-of-the-art results across diverse physics-related scenarios. Our project page is available at https://bwgzk-keke.github.io/DiffPhy/

📄 PDF Abstract BibTeX arXiv:2505.21653

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language ModelMultimodal Large Language ModelVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning

2026-04-21 · Haoyang Chen, Yi Liu, Jianzhi Shao, Tao Zhang 외 arxiv

Thinking LLMs produce reasoning traces before answering. Prior activation steering work mainly targets on shaping these traces. It remains less understood how answer tokens actually read and integrate the reasoning to pr…

Hume: Introducing System-2 Thinking in Visual-Language-Action Model

2025-05-27 · Haoming Song, Delin Qu, Yuanqi Yao, Qizhi Chen 외

Humans practice slow thinking before performing actual actions when handling complex tasks in the physical world. This thinking paradigm, recently, has achieved remarkable advancement in boosting Large Language Models (L…

DenoisingVision-Language-Action

Learning Physics-guided Face Relighting under Directional Light

2019-06-07 · CVPR 2020 6 · Thomas Nestmeyer, Jean-François Lalonde, Iain Matthews, Andreas M. Lehrmann

Relighting is an essential step in realistically transferring objects from a captured image into another environment. For example, authentic telepresence in Augmented Reality requires faces to be displayed and relit cons…

SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs

2025-10-06 · Dachuan Shi, Abedelkadir Asi, Keying Li, Xiangchi Yuan 외 arxiv

Recent work shows that, beyond discrete reasoning through explicit chain-of-thought steps, which are limited by the boundaries of natural languages, large language models (LLMs) can also reason continuously in latent spa…

Diffuse-CLoC: Guided Diffusion for Physics-based Character Look-ahead Control

2025-03-14 · Xiaoyu Huang, Takara Truong, Yunbo Zhang, Fangzhou Yu 외

We present Diffuse-CLoC, a guided diffusion framework for physics-based look-ahead control that enables intuitive, steerable, and physically realistic motion generation. While existing kinematics motion generation with d…

Action GenerationMotion Generationmotion in-betweening