paper-with-me

홈 › Papers

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation

2024-12-08 · CVPR 2025 1 · Shuwei Shi, Biao Gong, Xi Chen, Dandan Zheng, Shuai Tan, Zizheng Yang, Yuyuan Li, Jingwen He, Kecheng Zheng, Jingdong Chen, Ming Yang, Yinqiang Zheng

The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns, yet there lacks a reliable motion estimator for training such models on large-scale video set in the wild. Traditional metrics, e.g., SSIM or optical flow, are hard to generalize to arbitrary videos, while, it is very tough for human annotators to label the abstract motion intensity neither. Furthermore, the motion intensity shall reveal both local object motion and global camera movement, which has not been studied before. This paper addresses the challenge with a new motion estimator, capable of measuring the decoupled motion intensities of objects and cameras in video. We leverage the contrastive learning on randomly paired videos and distinguish the video with greater motion intensity. Such a paradigm is friendly for annotation and easy to scale up to achieve stable performance on motion estimation. We then present a new I2V model, named MotionStone, developed with the decoupled motion estimator. Experimental results demonstrate the stability of the proposed motion estimator and the state-of-the-art performance of MotionStone on I2V generation. These advantages warrant the decoupled motion estimator to serve as a general plug-in enhancer for both data processing and video generation training.

📄 PDF Abstract BibTeX arXiv:2412.05848

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningImage to Video GenerationMotion EstimationOptical Flow EstimationSSIMVideo Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

EMA: Effort Metric Attention for Anatomical Effort-Guided Human Motion Diffusion

2026-05-23 · Joshua Siy, Huakun Liu, Yutaro Hirao, Monica Perusquia-Hernandez 외 arxiv

Human motion diffusion models can synthesize action sequences from text, but controlling motion intensity remains challenging. Existing approaches rely on effort-related adverbs, which are ambiguous and fail to capture q…

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

2026-08-01 · Chenggong Hu, Shaoyin Ma, Yi Wang, Li Sun 외 arxiv

Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability and visual fidelity. Although implicit rep…

Talking Face GenerationContinuous Control

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion

2024-12-29 · Ashishkumar Gudmalwar, Ishan D. Biyani, Nirmesh Shah, Pankaj Wasnik 외

The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regulari…

Self-Supervised LearningVoice Conversion

Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion

2024-02-05 · Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma 외

Recent text-to-video diffusion models have achieved impressive progress. In practice, users often desire the ability to control object motion and camera movement independently for customized video creation. However, curr…

ObjectVideo Generation

Mojito: Motion Trajectory and Intensity Control for Video Generation

2024-12-12 · Xuehai He, Shuohang Wang, Jianwei Yang, Xiaoxia Wu 외

Recent advancements in diffusion models have shown great promise in producing high-quality video content. However, efficiently training diffusion models capable of integrating directional guidance and controllable motion…

Computational EfficiencyOptical Flow EstimationText-to-Video GenerationVideo Generation