paper-with-me

Papers

Inflation with Diffusion: Efficient Temporal Adaptation for Text-to-Video Super-Resolution

2024-01-18 · Xin Yuan, Jinoo Baek, Keyang Xu, Omer Tov, Hongliang Fei

We propose an efficient diffusion-based text-to-video super-resolution (SR) tuning approach that leverages the readily learned capacity of pixel level image diffusion model to capture spatial information for video generation. To accomplish this goal, we design an efficient architecture by inflating the weightings of the text-to-image SR model into our video generation framework. Additionally, we incorporate a temporal adapter to ensure temporal coherence across video frames. We investigate different tuning approaches based on our inflated architecture and report trade-offs between computational costs and super-resolution quality. Empirical evaluation, both quantitative and qualitative, on the Shutterstock video dataset, demonstrates that our approach is able to perform text-to-video SR generation with good visual quality and temporal consistency. To evaluate temporal coherence, we also present visualizations in video format in https://drive.google.com/drive/folders/1YVc-KMSJqOrEUdQWVaI-Yfu8Vsfu_1aO?usp=sharing .

📄 PDF Abstract BibTeX arXiv:2401.10404

Code (0)

등록된 구현이 없습니다.

Tasks

Super-ResolutionVideo GenerationVideo Super-Resolution

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Adapter 설명 없음

Similar Papers 제목 키워드 기반

Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation

2025-07-22 · Yaofang Liu, Yumeng Ren, Aitor Artola, Yuxuan Hu 외 arxiv

The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchronization of frame evolution imposed by conventional scalar timestep variabl…

Video Generation

FADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face Restoration

2026-07-07 · Jin Jiang, Jia Wang, Panwen Hu, Weiran Zhao 외 arxiv

Video face restoration (VFR) aims to recover high-quality and temporally consistent facial details from severely degraded video sequences; however, existing methods still struggle to balance spatial fidelity and temporal…

Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models

2024-02-22 · Yixuan Ren, Yang Zhou, Jimei Yang, Jing Shi 외

Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterp…

Video Generation

Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation

2023-06-13 · Shuai Yang, Yifan Zhou, Ziwei Liu, Chen Change Loy

Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains…

Patch MatchingTranslation

Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution

2024-03-25 · CVPR 2024 1 · Zhikai Chen, Fuchen Long, Zhaofan Qiu, Ting Yao 외

Diffusion models are just at a tipping point for image super-resolution task. Nevertheless, it is not trivial to capitalize on diffusion models for video super-resolution which necessitates not only the preservation of v…

DecoderDenoisingImage Super-ResolutionSuper-Resolution+3