paper-with-me

Papers

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment

2026-06-03 · Cong Wang, Hanxin Zhu, Jiayi Luo, Yonglin Tian, Xiaoqian Cheng, Peiyan Tu, Xin Jin, Long Chen, Zhibo Chen arxiv

Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincing. Nevertheless, the dynamics induced by pixel-level fitting do not naturally accommodate the regularities that govern real-world motion and interaction, resulting in persistent shortcomings in physical plausibility. To address this limitation, we propose \textbf{PILA} (Physics-Informed Latent Alignment), a framework that injects physics-structured latent guidance into the frozen flow-matching dynamics of pretrained video models. Specifically, PILA first employs anchored field estimation to map frozen-generator latents into an operational physical attribute bank organized by field-proxy slots, using observable motion as a kinematic anchor for constructing less directly observed proxies. To handle the heterogeneity of real-world dynamics, PILA adopts a mixture-of-experts design over physical categories. Label-prior masked expert routing selects category-specific operator experts, whose refinements are regularized by operational residuals abstracted from physical relations. Finally, the refined proxies are fused into the physical attribute bank and decoded into a correction to the flow-matching vector field, injecting physics-aware guidance while preserving the visual prior of the pretrained backbone. With staged adapter training on Wan 2.1-1.3B and direct transfer of the learned adapter to Wan 2.2-14B, PILA achieves state-of-the-art results on VBench-2.0, VideoPhy-2, and PhyGenBench in both visual quality and benchmark-measured physical plausibility.

📄 PDF Abstract BibTeX arXiv:2606.04737

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

2025-12-05 · Zijun Wang, Panwen Hu, Jing Wang, Terry Jingchen Zhang 외 arxiv

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-sca…

Video Generation

RoboScape: Physics-informed Embodied World Model

2025-06-29 · Yu Shang, Xin Zhang, Yinzhou Tang, Lei Jin 외

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current e…

3D geometryDepth EstimationDepth Predictionmodel+1

Unsupervised physics-informed disentanglement of multimodal data for high-throughput scientific discovery

2022-02-07 · Nathaniel Trask, Carianne Martinez, Kookjin Lee, Brad Boyce

We introduce physics-informed multimodal autoencoders (PIMA) - a variational inference framework for discovering shared information in multimodal scientific datasets representative of high-throughput testing. Individual …

DecoderDisentanglementscientific discoveryVariational Inference

Perception-Informed Neural Networks: Beyond Physics-Informed Neural Networks

2025-05-02 · Mehran Mazandarani, Marzieh Najariyan

This article introduces Perception-Informed Neural Networks (PrINNs), a framework designed to incorporate perception-based information into neural networks, addressing both systems with known and unknown physics laws or …

Mixture-of-Experts

WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation

2025-03-11 · Jing Wang, Ao Ma, Ke Cao, Jun Zheng 외

Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle to grasp abstract physical principles an…

Text-to-Video GenerationVideo Generation