paper-with-me

홈 › Papers

SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model

2024-01-01 · CVPR 2024 1 · Zhengang Li, Yan Kang, Yuchen Liu, Difan Liu, Tobias Hinz, Feng Liu, Yanzhi Wang

While AI-generated content has garnered significant attention achieving photo-realistic video synthesis remains a formidable challenge. Despite the promising advances in diffusion models for video generation quality the complex model architecture and substantial computational demands for both training and inference create a significant gap between these models and real-world applications. This paper presents SNED a superposition network architecture search method for efficient video diffusion model. Our method employs a supernet training paradigm that targets various model cost and resolution options using a weight-sharing method. Moreover we propose the supernet training sampling warm-up for fast training optimization. To showcase the flexibility of our method we conduct experiments involving both pixel-space and latent-space video diffusion models. The results demonstrate that our framework consistently produces comparable results across different model options with high efficiency. According to the experiment for the pixel-space video diffusion model we can achieve consistent video generation results simultaneously across 64 x 64 to 256 x 256 resolutions with a large range of model sizes from 640M to 1.6B number of parameters for pixel-space video diffusion models.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

On the Sum of Fisher-Snedecor F Variates and its Application to Maximal-Ratio Combining

2019-01-30

Capitalizing on the recently proposed Fisher-Snedecor F composite fading model, in this letter, we investigate the sum of independent but not identically distributed (i.n.i.d.) Fisher-Snedecor F variates. First, a novel …

Named Entity Disambiguation for Noisy Text

2017-06-28 · CONLL 2017 8 · Yotam Eshel, Noam Cohen, Kira Radinsky, Shaul Markovitch 외

We address the task of Named Entity Disambiguation (NED) for noisy text. We present WikilinksNED, a large-scale NED dataset of text fragments from the web, which is significantly noisier and more challenging than existin…

Entity DisambiguationEntity Embeddings

The Superposition of Diffusion Models Using the Itô Density Estimator

2024-12-23 · Marta Skreta, Lazar Atanackovic, Avishek Joey Bose, Alexander Tong 외

The Cambrian explosion of easily accessible pre-trained diffusion models suggests a demand for methods that combine multiple different pre-trained diffusion models without incurring the significant computational burden o…

Schrödinger's Bat: Diffusion Models Sometimes Generate Polysemous Words in Superposition

2022-11-23 · Jennifer C. White, Ryan Cotterell

Recent work has shown that despite their impressive capabilities, text-to-image diffusion models such as DALL-E 2 (Ramesh et al., 2022) can display strange behaviours when a prompt contains a word with multiple possible …

Video Diffusion Models

2022-04-07 · Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan 외

Generating temporally coherent high fidelity video is an important milestone in generative modeling research. We make progress towards this milestone by proposing a diffusion model for video generation that shows very pr…

Unconditional Video GenerationVideo GenerationVideo Prediction