paper-with-me

홈 › Papers

Orthogonal Spatial-temporal Distributional Transfer for 4D Generation

2026-03-05 · Wei Liu, Shengqiong Wu, Bobo Li, Haoyu Zhao, Hao Fei, Mong-Li Lee, Wynne Hsu arxiv

In the AIGC era, generating high-quality 4D content has garnered increasing research attention. Unfortunately, current 4D synthesis research is severely constrained by the lack of large-scale 4D datasets, preventing models from adequately learning the critical spatial-temporal features necessary for high-quality 4D generation, thus hindering progress in this domain. To combat this, we propose a novel framework that transfers rich spatial priors from existing 3D diffusion models and temporal priors from video diffusion models to enhance 4D synthesis. We develop a spatial-temporal-disentangled 4D (STD-4D) Diffusion model, which synthesizes 4D-aware videos through disentangled spatial and temporal latents. To facilitate the best feature transfer, we design a novel Orthogonal Spatial-temporal Distributional Transfer (Orster) mechanism, where the spatiotemporal feature distributions are carefully modeled and injected into the STD-4D Diffusion. Furthermore, during the 4D construction, we devise a spatial-temporal-aware HexPlane (ST-HexPlane) to integrate the transferred spatiotemporal features, thereby improving 4D deformation and 4D Gaussian feature modeling. Experiments demonstrate that our method significantly outperforms existing approaches, achieving superior spatial-temporal consistency and higher-quality 4D synthesis.

📄 PDF Abstract BibTeX arXiv:2603.05081

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Orthogonal Temporal Interpolation for Zero-Shot Video Recognition

2023-08-14 · Yan Zhu, Junbao Zhuo, Bin Ma, Jiajia Geng 외

Zero-shot video recognition (ZSVR) is a task that aims to recognize video categories that have not been seen during the model training process. Recently, vision-language models (VLMs) pre-trained on large-scale image-tex…

Video RecognitionZero-Shot Action RecognitionZero-Shot Action Recognition on HMDB51Zero-Shot Action Recognition on UCF101

Spatial Adapter: Structured Spatial Decomposition and Closed-Form Covariance for Frozen Predictors

2026-05-12 · Wen-Ting Wang, Wei-Ying Wu, Hao-Yun Huang, Xuan-Chun Wang arxiv

We present the Spatial Adapter, a parameter-efficient post-hoc layer that equips any frozen first-stage predictor with a structured spatial representation of its residual field and an induced closed-form spatial covarian…

DynaOD: Dynamic Origin-Destination Flow Generation with Discrete-to-Continuous Temporal Semantic Modeling

2026-06-08 · Jie Zhao, Xianqi Dai, Jie Feng, Huandong Wang 외 arxiv

Dynamic origin-destination (OD) flow generation seeks to synthesize realistic mobility dynamics from temporal context alone, without relying on historical OD observations. A key challenge is to translate semantic tempora…

CamMimic: Zero-Shot Image To Camera Motion Personalized Video Generation Using Diffusion Models

2025-04-13 · Pooja Guhan, Divya Kothandaraman, Tsung-Wei Huang, Guan-Ming Su 외

We introduce CamMimic, an innovative algorithm tailored for dynamic video editing needs. It is designed to seamlessly transfer the camera motion observed in a given reference video onto any scene of the user's choice in …

Video EditingVideo Generation

OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance

2026-03-19 · Cong Wang, Hanxin Zhu, Xiao Tang, Jiayi Luo 외 arxiv

Recent progress in video generation has led to substantial improvements in visual fidelity, yet ensuring physically consistent motion remains a fundamental challenge. Intuitively, this limitation can be attributed to the…

Video Generation