paper-with-me

Papers

VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning

2023-11-02 · Hong Chen, Xin Wang, Guanning Zeng, YiPeng Zhang, Yuwei Zhou, Feilin Han, Wenwu Zhu

Customized text-to-video generation aims to generate text-guided videos with customized user-given subjects, which has gained increasing attention recently. However, existing works are primarily limited to generating videos for a single subject, leaving the more challenging problem of customized multi-subject text-to-video generation largely unexplored. In this paper, we fill this gap and propose a novel VideoDreamer framework. VideoDreamer can generate temporally consistent text-guided videos that faithfully preserve the visual features of the given multiple subjects. Specifically, VideoDreamer leverages the pretrained Stable Diffusion with latent-code motion dynamics and temporal cross-frame attention as the base video generator. The video generator is further customized for the given multiple subjects by the proposed Disen-Mix Finetuning and Human-in-the-Loop Re-finetuning strategy, which can tackle the attribute binding problem of multi-subject generation. We also introduce MultiStudioBench, a benchmark for evaluating customized multi-subject text-to-video generation models. Extensive experiments demonstrate the remarkable ability of VideoDreamer to generate videos with new content such as new events and backgrounds, tailored to the customized multiple subjects. Our project page is available at https://videodreamer23.github.io/.

📄 PDF Abstract BibTeX arXiv:2311.00990

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeText-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
BASE 설명 없음

Similar Papers 제목 키워드 기반

DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control

2024-05-21 · Hong Chen, Xin Wang, YiPeng Zhang, Yuwei Zhou 외

Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and at…

AttributeMotion GenerationText-to-Video GenerationVideo Generation

GaussVideoDreamer: 3D Scene Generation with Video Diffusion and Inconsistency-Aware Gaussian Splatting

2025-04-14 · Junlin Hao, Peiheng Wang, Haoyang Wang, Xinggong Zhang 외

Single-image 3D scene reconstruction presents significant challenges due to its inherently ill-posed nature and limited input constraints. Recent advances have explored two promising directions: multiview generative mode…

3D Generation3D Scene ReconstructionOut-of-Distribution GeneralizationScene Generation+1

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

2025-05-07 · Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang 외

Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this p…

Human-Domain Subject-to-VideoSingle-Domain Subject-to-VideoVideo AlignmentVideo Generation

DreamVideo: Composing Your Dream Videos with Customized Subject and Motion

2023-12-07 · CVPR 2024 1 · Yujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan 외

Customized generation using diffusion models has made impressive progress in image generation, but remains unsatisfactory in the challenging video generation task, as it requires the controllability of both subjects and …

Image GenerationVideo Generation

MotionBooth: Motion-Aware Customized Text-to-Video Generation

2024-06-25 · Jianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang 외

In this work, we present MotionBooth, an innovative framework designed for animating customized subjects with precise control over both object and camera movements. By leveraging a few images of a specific object, we eff…

Text-to-Video GenerationVideo Generation