paper-with-me

홈 › Papers

DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution

2025-07-01 · Zhe Kong, Le Li, Yong Zhang, Feng Gao, Shaoshu Yang, Tao Wang, Kaihao Zhang, Zhuoliang Kang, Xiaoming Wei, Guanying Chen, Wenhan Luo arxiv

Real-world video super-resolution (VSR) presents significant challenges due to complex and unpredictable degradations. Although some recent methods utilize image diffusion models for VSR and have shown improved detail generation capabilities, they still struggle to produce temporally consistent frames. We attempt to use Stable Video Diffusion (SVD) combined with ControlNet to address this issue. However, due to the intrinsic image-animation characteristics of SVD, it is challenging to generate fine details using only low-quality videos. To tackle this problem, we propose DAM-VSR, an appearance and motion disentanglement framework for VSR. This framework disentangles VSR into appearance enhancement and motion control problems. Specifically, appearance enhancement is achieved through reference image super-resolution, while motion control is achieved through video ControlNet. This disentanglement fully leverages the generative prior of video diffusion models and the detail generation capabilities of image super-resolution models. Furthermore, equipped with the proposed motion-aligned bidirectional sampling strategy, DAM-VSR can conduct VSR on longer input videos. DAM-VSR achieves state-of-the-art performance on real-world data and AIGC data, demonstrating its powerful detail generation capabilities.

📄 PDF Abstract BibTeX arXiv:2507.01012

Code (0)

등록된 구현이 없습니다.

Tasks

Image Super-ResolutionVideo Super-Resolution

Similar Papers 제목 키워드 기반

IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head Generation

2024-12-05 · Sejong Yang, Seoung Wug Oh, Yang Zhou, Seon Joo Kim

We introduce a novel approach for high-resolution talking head generation from a single image and audio input. Prior methods using explicit face models, like 3D morphable models (3DMM) and facial landmarks, often fall sh…

DisentanglementTalking Head GenerationVideo Generation

MotionCrafter: One-Shot Motion Customization of Diffusion Models

2023-12-08 · Yuxin Zhang, Fan Tang, Nisha Huang, Haibin Huang 외

The essence of a video lies in its dynamic motions, including character actions, object movements, and camera movements. While text-to-video generative diffusion models have recently advanced in creating diverse contents…

DisentanglementMotion DisentanglementText-to-Video GenerationVideo Editing+1

LEO: Generative Latent Image Animator for Human Video Synthesis

2023-05-06 · Yaohui Wang, Xin Ma, Xinyuan Chen, Cunjian Chen 외

Spatio-temporal coherency is a major challenge in synthesizing high quality videos, particularly in synthesizing human videos that contain rich global and local deformations. To resolve this challenge, previous approache…

DisentanglementVideo Editing

FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control

2026-02-13 · Mingzhi Sheng, Zekai Gu, Peng Li, Cheng Lin 외 arxiv

Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental disentanglement of "appearance" and "mot…

Video Generation

PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement

2025-12-04 · Yu-Wei Zhan, Xin Wang, Hong Chen, Tongtong Feng 외 arxiv

Video Large Language Models (Video LLMs) have shown impressive performance across a wide range of video-language tasks. However, they often fail in scenarios requiring a deeper understanding of physical dynamics. This li…