paper-with-me

Papers

DeCo: Decoupled Human-Centered Diffusion Video Editing with Motion Consistency

2024-08-14 · Xiaojing Zhong, Xinyi Huang, Xiaofeng Yang, Guosheng Lin, Qingyao Wu

Diffusion models usher a new era of video editing, flexibly manipulating the video contents with text prompts. Despite the widespread application demand in editing human-centered videos, these models face significant challenges in handling complex objects like humans. In this paper, we introduce DeCo, a novel video editing framework specifically designed to treat humans and the background as separate editable targets, ensuring global spatial-temporal consistency by maintaining the coherence of each individual component. Specifically, we propose a decoupled dynamic human representation that utilizes a parametric human body prior to generate tailored humans while preserving the consistent motions as the original video. In addition, we consider the background as a layered atlas to apply text-guided image editing approaches on it. To further enhance the geometry and texture of humans during the optimization, we extend the calculation of score distillation sampling into normal space and image space. Moreover, we tackle inconsistent lighting between the edited targets by leveraging a lighting-aware video harmonizer, a problem previously overlooked in decompose-edit-combine approaches. Extensive qualitative and numerical experiments demonstrate that DeCo outperforms prior video editing methods in human-centered videos, especially in longer videos.

📄 PDF Abstract BibTeX arXiv:2408.07481

Code (0)

등록된 구현이 없습니다.

Tasks

text-guided-image-editingVideo Editing

Similar Papers 제목 키워드 기반

Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model

2024-04-02 · CVPR 2024 1 · Xu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin 외

Co-speech gestures, if presented in the lively form of videos, can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons, resulting in the omission …

Video Generation

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation

2024-12-08 · CVPR 2025 1 · Shuwei Shi, Biao Gong, Xi Chen, Dandan Zheng 외

The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate divers…

Contrastive LearningImage to Video GenerationMotion EstimationOptical Flow Estimation+2

Decoupled Video Generation with Chain of Training-free Diffusion Model Experts

2024-08-24 · Wenhao Li, Yichao Cao, Xiu Su, Xi Lin 외

Video generation models hold substantial potential in areas such as filmmaking. However, current video diffusion models need high computational costs and produce suboptimal results due to extreme complexity of video gene…

DenoisingVideo Generation

Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning

2025-06-05 · Yue Ma, Yulong Liu, Qiyuan Zhu, Ayden Yang 외

Recently, breakthroughs in the video diffusion transformer have shown remarkable capabilities in diverse motion generations. As for the motion-transfer task, current methods mainly use two-stage Low-Rank Adaptations (LoR…

TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation

2025-04-11 · CVPR 2025 1 · Ruineng Li, Daitao Xing, Huiming Sun, Yuanzhou Ha 외

Human-centric motion control in video generation remains a critical challenge, particularly when jointly controlling camera movements and human poses in scenarios like the iconic Grammy Glambot moment. While recent video…

DisentanglementVideo Generation