paper-with-me

홈 › Papers

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation

2024-11-30 · CVPR 2025 1 · Qiyao Xue, Xiangyu Yin, Boyuan Yang, Wei Gao

Text-to-video (T2V) generation has been recently enabled by transformer-based diffusion models, but current T2V models lack capabilities in adhering to the real-world common knowledge and physical rules, due to their limited understanding of physical realism and deficiency in temporal modeling. Existing solutions are either data-driven or require extra model inputs, but cannot be generalizable to out-of-distribution domains. In this paper, we present PhyT2V, a new data-independent T2V technique that expands the current T2V model's capability of video generation to out-of-distribution domains, by enabling chain-of-thought and step-back reasoning in T2V prompting. Our experiments show that PhyT2V improves existing T2V models' adherence to real-world physical rules by 2.3x, and achieves 35% improvement compared to T2V prompt enhancers. The source codes are available at: https://github.com/pittisl/PhyT2V.

📄 PDF Abstract BibTeX arXiv:2412.00596

Code (1)

pittisl/phyt2v 공식 구현 pytorch

Tasks

Text-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Phythesis: Physics-Guided Evolutionary Scene Synthesis for Energy-Efficient Data Center Design via LLMs

2025-12-11 · Minghao LI, Ruihang Wang, Rui Tan, Yonggang Wen arxiv

Data center (DC) infrastructure serves as the backbone to support the escalating demand for computing capacity. Traditional design methodologies that blend human expertise with specialized simulation tools scale poorly w…

Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement

2025-11-25 · Yang Liu, Xilin Zhao, Peisong Wen, Siran Dai 외 arxiv

Recent progress in video generation has led to impressive visual quality, yet current models still struggle to produce results that align with real-world physical principles. To this end, we propose an iterative self-ref…

Video Generation

PhyTracker: An Online Tracker for Phytoplankton

2024-06-29 · Yang Yu, Qingxuan Lv, Yuezun Li, Zhiqiang Wei 외

Phytoplankton, a crucial component of aquatic ecosystems, requires efficient monitoring to understand marine ecological processes and environmental conditions. Traditional phytoplankton monitoring methods, relying on non…

VeREFINE: Integrating Object Pose Verification with Physics-guided Iterative Refinement

2019-09-12 · Dominik Bauer, Timothy Patten, Markus Vincze

Accurate and robust object pose estimation for robotics applications requires verification and refinement steps. In this work, we propose to integrate hypotheses verification with object pose refinement guided by physics…

ObjectPose Estimation

RASPRef: Retrieval-Augmented Self-Supervised Prompt Refinement for Large Reasoning Models

2026-03-27 · Rahul Soni arxiv

Recent reasoning-focused language models such as DeepSeek R1 and OpenAI o1 have demonstrated strong performance on structured reasoning benchmarks including GSM8K, MATH, and multi-hop question answering tasks. However, t…

Multi-hop Question AnsweringMathematical Reasoning