paper-with-me

Papers

Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

2023-11-17 · Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, Ishan Misra

We present Emu Video, a text-to-video generation model that factorizes the generation into two steps: first generating an image conditioned on the text, and then generating a video conditioned on the text and the generated image. We identify critical design decisions--adjusted noise schedules for diffusion, and multi-stage training that enable us to directly generate high quality and high resolution videos, without requiring a deep cascade of models as in prior work. In human evaluations, our generated videos are strongly preferred in quality compared to all prior work--81% vs. Google's Imagen Video, 90% vs. Nvidia's PYOCO, and 96% vs. Meta's Make-A-Video. Our model outperforms commercial solutions such as RunwayML's Gen2 and Pika Labs. Finally, our factorizing approach naturally lends itself to animating images based on a user's text prompt, where our generations are preferred 96% over prior work.

📄 PDF Abstract BibTeX arXiv:2311.10709

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video GenerationVideo Generation

Similar Papers 제목 키워드 기반

DeRA: Decoupled Representation Alignment for Video Tokenization

2025-12-04 · Pengbo Guo, Junke Wang, Zhen Xing, Chengxu Liu 외 arxiv

This paper presents DeRA, a novel 1D video tokenizer that decouples the spatial-temporal representation learning in video tokenization to achieve better training efficiency and performance. Specifically, DeRA maintains a…

Representation LearningVideo Generation

Subject-driven Video Generation via Disentangled Identity and Motion

2025-04-23 · Daneul Kim, Jingxu Zhang, Wonjoon Jin, Sunghyun Cho 외

We propose to train a subject-driven customized video generation model through decoupling the subject-specific learning from temporal dynamics in zero-shot without additional tuning. A traditional method for video custom…

Subject-driven Video GenerationVideo Generation

Context-aware Talking Face Video Generation

2024-02-28 · Meidai Xuanyuan, Yuwang Wang, Honglei Guo, Qionghai Dai

In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroun…

Video GenerationVideo Synchronization

Scaling Zero-Shot Reference-to-Video Generation

2025-12-07 · Zijian Zhou, Shikun Liu, Haozhe Liu, Haonan Qiu 외 arxiv

Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V methods are hindered by the reliance on explic…

Video Generation

DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos

2024-05-03 · Wen-Hsuan Chu, Lei Ke, Katerina Fragkiadaki

View-predictive generative models provide strong priors for lifting object-centric images and videos into 3D and 4D through rendering and score distillation objectives. A question then remains: what about lifting complet…

Depth EstimationDepth PredictionNovel View SynthesisObject+2