paper-with-me

Papers

MotionBooth: Motion-Aware Customized Text-to-Video Generation

2024-06-25 · Jianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang, Qianyu Zhou, Yining Li, Yunhai Tong, Kai Chen

In this work, we present MotionBooth, an innovative framework designed for animating customized subjects with precise control over both object and camera movements. By leveraging a few images of a specific object, we efficiently fine-tune a text-to-video model to capture the object's shape and attributes accurately. Our approach presents subject region loss and video preservation loss to enhance the subject's learning performance, along with a subject token cross-attention loss to integrate the customized subject with motion control signals. Additionally, we propose training-free techniques for managing subject and camera motions during inference. In particular, we utilize cross-attention map manipulation to govern subject motion and introduce a novel latent shift module for camera movement control as well. MotionBooth excels in preserving the appearance of subjects while simultaneously controlling the motions in generated videos. Extensive quantitative and qualitative evaluations demonstrate the superiority and effectiveness of our method. Our project page is at https://jianzongwu.github.io/projects/motionbooth

📄 PDF Abstract BibTeX arXiv:2406.17758

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video GenerationVideo Generation

Similar Papers 제목 키워드 기반

Still-Moving: Customized Video Generation without Customized Video Data

2024-07-11 · Hila Chefer, Shiran Zada, Roni Paiss, Ariel Ephrat 외

Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation i…

Video Generation

MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models

2024-12-02 · Xiaomin Li, Xu Jia, Qinghe Wang, Haiwen Diao 외

Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exhibit significant limitations when genera…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

DreamVideo: Composing Your Dream Videos with Customized Subject and Motion

2023-12-07 · CVPR 2024 1 · Yujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan 외

Customized generation using diffusion models has made impressive progress in image generation, but remains unsatisfactory in the challenging video generation task, as it requires the controllability of both subjects and …

Image GenerationVideo Generation

CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training

2024-12-20 · Xiuli Bi, Jian Lu, Bo Liu, Xiaodong Cun 외

Benefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the p…

parameter-efficient fine-tuningVideo Generation

Bring Your Dreams to Life: Continual Text-to-Video Customization

2025-12-05 · Jiahua Dong, Xudong Wang, Wenqi Liang, Zongyan Han 외 arxiv

Customized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that personalized concepts remain static and do no…

Text-to-Video GenerationNoise Estimation