paper-with-me

홈 › Papers

HARIVO: Harnessing Text-to-Image Models for Video Generation

2024-10-10 · Mingi Kwon, Seoung Wug Oh, Yang Zhou, Difan Liu, Joon-Young Lee, Haoran Cai, Baqiao Liu, Feng Liu, Youngjung Uh

We present a method to create diffusion-based video models from pretrained Text-to-Image (T2I) models. Recently, AnimateDiff proposed freezing the T2I model while only training temporal layers. We advance this method by proposing a unique architecture, incorporating a mapping network and frame-wise tokens, tailored for video generation while maintaining the diversity and creativity of the original T2I model. Key innovations include novel loss functions for temporal smoothness and a mitigating gradient sampling technique, ensuring realistic and temporally consistent video generation despite limited public video data. We have successfully integrated video-specific inductive biases into the architecture and loss functions. Our method, built on the frozen StableDiffusion model, simplifies training processes and allows for seamless integration with off-the-shelf models like ControlNet and DreamBooth. project page: https://kwonminki.github.io/HARIVO

📄 PDF Abstract BibTeX arXiv:2410.07763

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityVideo Generation

Similar Papers 제목 키워드 기반

UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

2024-02-20 · Jianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo 외

Recent advances in text-guided video editing have showcased promising results in appearance editing (e.g., stylization). However, video motion editing in the temporal dimension (e.g., from eating to waving), which distin…

Video Editing

SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation

2025-08-01 · Kien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen 외 arxiv

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focu…

Video Generation

Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation

2024-05-24 · Shentong Mo, Yapeng Tian

In recent developments, the Mamba architecture, known for its selective state space approach, has shown potential in the efficient modeling of long sequences. However, its application in image generation remains underexp…

Image GenerationMambaVideo Generation

\textsc{Gen2Real}: Towards Demo-Free Dexterous Manipulation by Harnessing Generated Video

2025-09-16 · Kai Ye, Yuhang Wu, Shuyuan Hu, Junliang Li 외 arxiv

Dexterous manipulation remains a challenging robotics problem, largely due to the difficulty of collecting extensive human demonstrations for learning. In this paper, we introduce \textsc{Gen2Real}, which replaces costly…

Depth EstimationVideo Generation

ReCamMaster: Camera-Controlled Generative Rendering from A Single Video

2025-03-14 · Jianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang 외

Camera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video cr…

Super-ResolutionVideo GenerationVideo Stabilization