Mojito: Motion Trajectory and Intensity Control for Video Generation
Recent advancements in diffusion models have shown great promise in producing high-quality video content. However, efficiently training diffusion models capable of integrating directional guidance and controllable motion intensity remains a challenging and under-explored area. This paper introduces Mojito, a diffusion model that incorporates both \textbf{Mo}tion tra\textbf{j}ectory and \textbf{i}ntensi\textbf{t}y contr\textbf{o}l for text to video generation. Specifically, Mojito features a Directional Motion Control module that leverages cross-attention to efficiently direct the generated object's motion without additional training, alongside a Motion Intensity Modulator that uses optical flow maps generated from videos to guide varying levels of motion intensity. Extensive experiments demonstrate Mojito's effectiveness in achieving precise trajectory and intensity control with high computational efficiency, generating motion patterns that closely match specified directions and intensities, providing realistic dynamics that align well with natural motion in real-world scenarios.
Code (0)
등록된 구현이 없습니다.
Tasks
Computational EfficiencyOptical Flow EstimationText-to-Video GenerationVideo GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downstream planner predicts trajectories cond…
Instruction FollowingAutonomous DrivingMojito: LLM-Aided Motion Instructor with Jitter-Reduced Inertial Tokens
Human bodily movements convey critical insights into action intentions and cognitive processes, yet existing multimodal systems primarily focused on understanding human motion via language, vision, and audio, which strug…
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
Recent advances in trajectory-controllable video generation have achieved remarkable progress. Previous methods mainly use adapter-based architectures for precise motion control along predefined trajectories. However, al…
Video GenerationMotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate divers…
Contrastive LearningImage to Video GenerationMotion EstimationOptical Flow Estimation+2MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control th…
Image to Video GenerationObjectVideo Generation