paper-with-me

홈 › Papers

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

2023-10-30 · Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, Chao Weng, Ying Shan

Video generation has increasingly gained interest in both academia and industry. Although commercial tools can generate plausible videos, there is a limited number of open-source models available for researchers and engineers. In this work, we introduce two diffusion models for high-quality video generation, namely text-to-video (T2V) and image-to-video (I2V) models. T2V models synthesize a video based on a given text input, while I2V models incorporate an additional image input. Our proposed T2V model can generate realistic and cinematic-quality videos with a resolution of $1024 \times 576$, outperforming other open-source T2V models in terms of quality. The I2V model is designed to produce videos that strictly adhere to the content of the provided reference image, preserving its content, structure, and style. This model is the first open-source I2V foundation model capable of transforming a given image into a video clip while maintaining content preservation constraints. We believe that these open-source video generation models will contribute significantly to the technological advancements within the community.

📄 PDF Abstract BibTeX arXiv:2310.19512

Code (3)

ailab-cvc/videocrafter 공식 구현 pytorch
invictus717/interactivevideo pytorch
videocrafter/videocrafter pytorch

Tasks

Text-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

VEnhancer: Generative Space-Time Enhancement for Video Generation

2024-07-10 · Jingwen He, Tianfan Xue, Dongyang Liu, Xinqi Lin 외

We present VEnhancer, a generative space-time enhancement framework that improves the existing text-to-video results by adding more details in spatial domain and synthetic detailed motion in temporal domain. Given a gene…

Data AugmentationSuper-ResolutionVideo GenerationVideo Super-Resolution

VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

2024-01-17 · CVPR 2024 1 · Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia 외

Text-to-video generation aims to produce a video based on a given prompt. Recently, several commercial video models have been able to generate plausible videos with minimal noise, excellent details, and high aesthetic sc…

Text-to-Video GenerationVideo Generation

V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models

2025-08-05 · Jisoo Kim, Wooseok Seo, Junwan Kim, Seungho Park 외 arxiv

With growing interest in deploying text-to-video (T2V) models in resource-constrained environments, reducing their high computational cost has become crucial, leading to extensive research on pruning and knowledge distil…

Knowledge DistillationVideo Generation

Does Semantic Noise Initialization Transfer from Images to Videos? A Paired Diagnostic Study

2026-03-03 · Yixiao Jing, Chaoyu Zhang, Zixuan Zhong, Peizhou Huang arxiv

Semantic noise initialization has been reported to improve robustness and controllability in image diffusion models. Whether these gains transfer to text-to-video (T2V) generation remains unclear, since temporal coupling…

Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization

2025-04-11 · Jialu Li, Shoubin Yu, Han Lin, Jaemin Cho 외

Recent advancements in text-to-video (T2V) diffusion models have significantly enhanced the visual quality of the generated videos. However, even recent T2V models find it challenging to follow text descriptions accurate…

DenoisingObjectSemantic SegmentationText-to-Video Generation+1