paper-with-me

Papers

RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution

2025-07-25 · Weisong Zhao, Jingkai Zhou, Xiangyu Zhu, Weihua Chen, Xiao-Yu Zhang, Zhen Lei, Fan Wang arxiv

Video Super-Resolution (VSR) has achieved significant progress through diffusion models, effectively addressing the over-smoothing issues inherent in GAN-based methods. Despite recent advances, three critical challenges persist in VSR community: 1) Inconsistent modeling of temporal dynamics in foundational models; 2) limited high-frequency detail recovery under complex real-world degradations; and 3) insufficient evaluation of detail enhancement and 4K super-resolution, as current methods primarily rely on 720P datasets with inadequate details. To address these challenges, we propose RealisVSR, a high-frequency detail-enhanced video diffusion model with three core innovations: 1) Consistency Preserved ControlNet (CPC) architecture integrated with the Wan2.1 video diffusion to model the smooth and complex motions and suppress artifacts; 2) High-Frequency Rectified Diffusion Loss (HR-Loss) combining wavelet decomposition and HOG feature constraints for texture restoration; 3) RealisVideo-4K, the first public 4K VSR benchmark containing 1,000 high-definition video-text pairs. Leveraging the advanced spatio-temporal guidance of Wan2.1, our method requires only 5-25% of the training data volume compared to existing approaches. Extensive experiments on VSR benchmarks (REDS, SPMCS, UDM10, YouTube-HQ, VideoLQ, RealisVideo-720P) demonstrate our superiority, particularly in ultra-high-resolution scenarios.

📄 PDF Abstract BibTeX arXiv:2507.19138

Code (0)

등록된 구현이 없습니다.

Tasks

Video Super-Resolution

Similar Papers 제목 키워드 기반

FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation

2024-12-01 · Yunpeng Bai, QiXing Huang

Monocular Depth Estimation (MDE) is essential for applications like 3D scene reconstruction, autonomous navigation, and AI content creation. However, robust MDE remains challenging due to noisy real-world data and distri…

3D Scene ReconstructionAutonomous NavigationDepth EstimationMonocular Depth Estimation

Generative Detail Enhancement for Physically Based Materials

2025-02-19 · Saeed Hadadan, Benedikt Bitterli, Tizian Zeltner, Jan Novák 외

We present a tool for enhancing the detail of physically based materials using an off-the-shelf diffusion model and inverse rendering. Our goal is to enhance the visual fidelity of materials with detail that is often ted…

Inverse Rendering

MPOD123: One Image to 3D Content Generation Using Mask-enhanced Progressive Outline-to-Detail Optimization

2024-01-01 · CVPR 2024 1 · Jimin Xu, Tianbao Wang, Tao Jin, Shengyu Zhang 외

Recent advancements in single image driven 3D content generation have been propelled by leveraging prior knowledge from pretrained 2D diffusion models. However the 3D content generated by existing methods often exhib…

Image to 3D

Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration

2024-12-01 · Haoze Sun, Wenbo Li, Jiayue Liu, Kaiwen Zhou 외

Generalization has long been a central challenge in real-world image restoration. While recent diffusion-based restoration methods, which leverage generative priors from text-to-image models, have made progress in recove…

Image Restoration

DiffVSR: Enhancing Real-World Video Super-Resolution with Diffusion Models for Advanced Visual Quality and Temporal Consistency

2025-01-17 · Xiaohui Li, Yihao Liu, Shuo Cao, Ziyan Chen 외

Diffusion models have demonstrated exceptional capabilities in image generation and restoration, yet their application to video super-resolution faces significant challenges in maintaining both high fidelity and temporal…

DecoderImage GenerationSuper-ResolutionVideo Super-Resolution