paper-with-me

홈 › Papers

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation

2026-06-25 · Kexu Cheng, Zicheng Liu, Mingju Gao, Chunhe Song, Hao Tang arxiv

Developing physically aware video generation models remains a significant challenge due to the difficulty in capturing diverse physical phenomena, such as thermal dynamics, mechanics, and optics. In this work, we introduce PhysRAG, a novel pipeline that enhances physical awareness in video generation through Retrieval-Augmented Generation (RAG). To address the issue of limited high-quality data, we design a two-stage data filtering pipeline based on the WISA-80K dataset, resulting in a curated set of 7K high-quality videos for training. Furthermore, we construct a physical video database and develop a mechanism to inject physical knowledge into a video diffusion model using learnable queries. Our method achieves state-of-the-art performance in both visual quality and physical rule compliance, surpassing existing models in benchmarks such as PhyGenBench and VBench. We conduct extensive ablation studies to validate the effectiveness of our key components, including the data filtering pipeline, RAG mechanism, and method for physical information extraction. To facilitate future research, our code, data, and models are prepared for release at https://github.com/sediment1024/PhysRAG.

📄 PDF Abstract BibTeX arXiv:2606.26916

Code (0)

등록된 구현이 없습니다.

Tasks

Information ExtractionVideo Generation

Similar Papers 제목 키워드 기반

WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation

2025-03-11 · Jing Wang, Ao Ma, Ke Cao, Jun Zheng 외

Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle to grasp abstract physical principles an…

Text-to-Video GenerationVideo Generation

PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning

2025-10-15 · Sihui Ji, Xi Chen, Xin Tao, Pengfei Wan 외 arxiv

Video generation models nowadays are capable of generating visually realistic videos, but often fail to adhere to physical laws, limiting their ability to generate physically plausible videos and serve as ''world models'…

Representation LearningReinforcement LearningVideo Generation

Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement

2025-11-25 · Yang Liu, Xilin Zhao, Peisong Wen, Siran Dai 외 arxiv

Recent progress in video generation has led to impressive visual quality, yet current models still struggle to produce results that align with real-world physical principles. To this end, we propose an iterative self-ref…

Video Generation

RoboScape: Physics-informed Embodied World Model

2025-06-29 · Yu Shang, Xin Zhang, Yinzhou Tang, Lei Jin 외

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current e…

3D geometryDepth EstimationDepth Predictionmodel+1

Enhancing Scene Transition Awareness in Video Generation via Post-Training

2025-07-24 · Hanwen Shen, Jiajie Lu, Yupeng Cao, Xiaonan Yang arxiv

Recent advances in AI-generated video have shown strong performance on \emph{text-to-video} tasks, particularly for short clips depicting a single scene. However, current models struggle to generate longer videos with co…

Scene GenerationVideo Generation