paper-with-me

Papers

Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning

2024-10-31 · Penghui Ruan, Pichao Wang, Divya Saxena, Jiannong Cao, Yuhui Shi

Despite advancements in Text-to-Video (T2V) generation, producing videos with realistic motion remains challenging. Current models often yield static or minimally dynamic outputs, failing to capture complex motions described by text. This issue stems from the internal biases in text encoding, which overlooks motions, and inadequate conditioning mechanisms in T2V generation models. To address this, we propose a novel framework called DEcomposed MOtion (DEMO), which enhances motion synthesis in T2V generation by decomposing both text encoding and conditioning into content and motion components. Our method includes a content encoder for static elements and a motion encoder for temporal dynamics, alongside separate content and motion conditioning mechanisms. Crucially, we introduce text-motion and video-motion supervision to improve the model's understanding and generation of motion. Evaluations on benchmarks such as MSR-VTT, UCF-101, WebVid-10M, EvalCrafter, and VBench demonstrate DEMO's superior ability to produce videos with enhanced motion dynamics while maintaining high visual quality. Our approach significantly advances T2V generation by integrating comprehensive motion understanding directly from textual descriptions. Project page: https://PR-Ryan.github.io/DEMO-project/

📄 PDF Abstract BibTeX arXiv:2410.24219

Code (1)

pr-ryan/demo 공식 구현 pytorch

Tasks

Motion SynthesisText-to-Video GenerationVideo Generation

Similar Papers 제목 키워드 기반

LaMD: Latent Motion Diffusion for Image-Conditional Video Generation

2023-04-23 · Yaosi Hu, Zhenzhong Chen, Chong Luo

The video generation field has witnessed rapid improvements with the introduction of recent diffusion models. While these models have successfully enhanced appearance quality, they still face challenges in generating coh…

Motion GenerationVideo GenerationVideo Reconstruction

RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space

2025-08-12 · Jingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao 외 arxiv

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground s…

Video Generation

Text2Performer: Text-Driven Human Video Generation

2023-04-17 · ICCV 2023 1 · Yuming Jiang, Shuai Yang, Tong Liang Koh, Wayne Wu 외

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts des…

Video Generation

MoCoGAN: Decomposing Motion and Content for Video Generation

2017-07-17 · CVPR 2018 6 · Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, Jan Kautz

Visual signals in a video can be divided into content and motion. While content specifies which objects are in the video, motion describes their dynamics. Based on this prior, we propose the Motion and Content decomposed…

Generative Adversarial NetworkVideo Generation

GenSpan: Generation-Calibrated Motion Span Priors for Multi-Verb Video Corpus Moment Retrieval

2026-03-23 · Yunzhuo Sun, Xinyue Liu, Yanyang Li, Nanding Wu 외 arxiv

Video Corpus Moment Retrieval (VCMR) aims to retrieve both the correct video and its temporal segment corresponding to a natural-language query, a task that is especially challenging for multi-verb queries where temporal…

Moment Retrieval