paper-with-me

홈 › Papers

Mojito: Motion Trajectory and Intensity Control for Video Generation

2024-12-12 · Xuehai He, Shuohang Wang, Jianwei Yang, Xiaoxia Wu, Yiping Wang, Kuan Wang, Zheng Zhan, Olatunji Ruwase, Yelong Shen, Xin Eric Wang

Recent advancements in diffusion models have shown great promise in producing high-quality video content. However, efficiently training diffusion models capable of integrating directional guidance and controllable motion intensity remains a challenging and under-explored area. This paper introduces Mojito, a diffusion model that incorporates both \textbf{Mo}tion tra\textbf{j}ectory and \textbf{i}ntensi\textbf{t}y contr\textbf{o}l for text to video generation. Specifically, Mojito features a Directional Motion Control module that leverages cross-attention to efficiently direct the generated object's motion without additional training, alongside a Motion Intensity Modulator that uses optical flow maps generated from videos to guide varying levels of motion intensity. Extensive experiments demonstrate Mojito's effectiveness in achieving precise trajectory and intensity control with high computational efficiency, generating motion patterns that closely match specified directions and intensities, providing realistic dynamics that align well with natural motion in real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2412.08948

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyOptical Flow EstimationText-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

2026-07-26 · Zhijing Cheng, Xuancheng Zhang, Donglin Di, Lei Fan 외 arxiv

End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downstream planner predicts trajectories cond…

Instruction FollowingAutonomous Driving

Mojito: LLM-Aided Motion Instructor with Jitter-Reduced Inertial Tokens

2025-02-22 · Ziwei Shan, Yaoyu He, Chengfeng Zhao, Jiashen Du 외

Human bodily movements convey critical insights into action intentions and cognitive processes, yet existing multimodal systems primarily focused on understanding human motion via language, vision, and audio, which strug…

FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance

2026-03-12 · Quanhao Li, Zhen Xing, Rui Wang, Haidong Cao 외 arxiv

Recent advances in trajectory-controllable video generation have achieved remarkable progress. Previous methods mainly use adapter-based architectures for precise motion control along predefined trajectories. However, al…

Video Generation

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation

2024-12-08 · CVPR 2025 1 · Shuwei Shi, Biao Gong, Xi Chen, Dandan Zheng 외

The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate divers…

Contrastive LearningImage to Video GenerationMotion EstimationOptical Flow Estimation+2

MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance

2025-03-20 · Quanhao Li, Zhen Xing, Rui Wang, HUI ZHANG 외

Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control th…

Image to Video GenerationObjectVideo Generation