paper-with-me

홈 › Papers

Motion Control for Enhanced Complex Action Video Generation

2024-11-13 · Qiang Zhou, Shaofeng Zhang, Nianzu Yang, Ye Qian, Hao Li

Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to precisely convey intricate motion details. To address this, we propose a novel framework, MVideo, designed to produce long-duration videos with precise, fluid actions. MVideo overcomes the limitations of text prompts by incorporating mask sequences as an additional motion condition input, providing a clearer, more accurate representation of intended actions. Leveraging foundational vision models such as GroundingDINO and SAM2, MVideo automatically generates mask sequences, enhancing both efficiency and robustness. Our results demonstrate that, after training, MVideo effectively aligns text prompts with motion conditions to produce videos that simultaneously meet both criteria. This dual control mechanism allows for more dynamic video generation by enabling alterations to either the text prompt or motion condition independently, or both in tandem. Furthermore, MVideo supports motion condition editing and composition, facilitating the generation of videos with more complex actions. MVideo thus advances T2V motion generation, setting a strong benchmark for improved action depiction in current video diffusion models. Our project page is available at https://mvideo-v1.github.io/.

📄 PDF Abstract BibTeX arXiv:2411.08328

Code (0)

등록된 구현이 없습니다.

Tasks

Motion GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Toward Rich Video Human-Motion2D Generation

2025-06-17 · Ruihao Xi, Xuekuan Wang, Yongcheng Li, Shuhua Li 외

Generating realistic and controllable human motions, particularly those involving rich multi-character interactions, remains a significant challenge due to data scarcity and the complexities of modeling inter-personal dy…

Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation

2025-04-21 · Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang 외

Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with high-quality annotations for both aspects. …

Video Generation

AKiRa: Augmentation Kit on Rays for optical video generation

2024-12-18 · CVPR 2025 1 · Xi Wang, Robin Courant, Marc Christie, Vicky Kalogeiton

Recent advances in text-conditioned video diffusion have greatly improved video quality. However, these methods offer limited or sometimes no control to users on camera aspects, including dynamic camera motion, zoom, dis…

Video Generation

Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

2025-02-20 · Pengxiang Ding, Jianfei Ma, Xinyang Tong, Binghong Zou 외

This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to data scarcity. We propose Humanoid-VLA, a…

Data AugmentationHumanoid ControlMotion Generation

ReVision: High-Quality, Low-Cost Video Generation with Explicit 3D Physics Modeling for Complex Motion and Interaction

2025-04-30 · Qihao Liu, Ju He, Qihang Yu, Liang-Chieh Chen 외

In recent years, video generation has seen significant advancements. However, challenges still persist in generating complex motions and interactions. To address these challenges, we introduce ReVision, a plug-and-play f…

Video Generation