paper-with-me

Papers

Moaw: Unleashing Motion Awareness for Video Diffusion Models

2026-01-19 · Tianqi Zhang, Ziyi Wang, Wenzhao Zheng, Weiliang Chen, Yuanhui Huang, Zhengyang Huang, Jie Zhou, Jiwen Lu arxiv

Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks such as optical flow prediction and tracking in a zero-shot setting. Motivated by these findings, we investigate whether supervised training can more fully harness the tracking capability of video diffusion models. To this end, we propose Moaw, a framework that unleashes motion awareness for video diffusion models and leverages it to facilitate motion transfer. Specifically, we train a diffusion model for motion perception, shifting its modality from image-to-video generation to video-to-dense-tracking. We then construct a motion-labeled dataset to identify features that encode the strongest motion information, and inject them into a structurally identical video generation model. Owing to the homogeneity between the two networks, these features can be naturally adapted in a zero-shot manner, enabling motion transfer without additional adapters. Our work provides a new paradigm for bridging generative modeling and motion understanding, paving the way for more unified and controllable video learning frameworks.

📄 PDF Abstract BibTeX arXiv:2601.12761

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

DanceGRPO: Unleashing GRPO on Visual Generation

2025-05-12 · Zeyue Xue, Jie Wu, Yu Gao, Fangyuan Kong 외

Recent breakthroughs in generative models-particularly diffusion models and rectified flows-have revolutionized visual content creation, yet aligning model outputs with human preferences remains a critical challenge. Exi…

Denoisingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness

2024-06-11 · Zihui Xue, Mi Luo, Changan Chen, Kristen Grauman

We study the problem of precisely swapping objects in videos, with a focus on those interacted with by hands, given one user-provided reference object image. Despite the great advancements that diffusion models have made…

ObjectVideo EditingVideo Generation

CPA: Camera-pose-awareness Diffusion Transformer for Video Generation

2024-12-02 · Yuelei Wang, Jian Zhang, PengTao Jiang, Hao Zhang 외

Despite the significant advancements made by Diffusion Transformer (DiT)-based methods in video generation, there remains a notable gap with controllable camera pose perspectives. Existing works such as OpenSora do NOT a…

Text-to-Video GenerationVideo Generation

Articulate That Object Part (ATOP): 3D Part Articulation from Text and Motion Personalization

2025-02-11 · Aditya Vora, Sauradip Nag, Hao Zhang

We present ATOP (Articulate That Object Part), a novel method based on motion personalization to articulate a 3D object with respect to a part and its motion as prescribed in a text prompt. Specifically, the text input a…

Image GenerationMotion GenerationObjectVideo Generation

Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach

2025-02-05 · Yunuo Chen, Junli Cao, Anil Kag, Vidit Goel 외

We present a novel video generation framework that integrates 3-dimensional geometry and dynamic awareness. To achieve this, we augment 2D videos with 3D point trajectories and align them in pixel space. The resulting 3D…

Video Generation