paper-with-me

Papers

Learning Spatiotemporal Frequency-Transformer for Low-Quality Video Super-Resolution

2022-12-27 · Zhongwei Qiu, Huan Yang, Jianlong Fu, Daochang Liu, Chang Xu, Dongmei Fu

Video Super-Resolution (VSR) aims to restore high-resolution (HR) videos from low-resolution (LR) videos. Existing VSR techniques usually recover HR frames by extracting pertinent textures from nearby frames with known degradation processes. Despite significant progress, grand challenges are remained to effectively extract and transmit high-quality textures from high-degraded low-quality sequences, such as blur, additive noises, and compression artifacts. In this work, a novel Frequency-Transformer (FTVSR) is proposed for handling low-quality videos that carry out self-attention in a combined space-time-frequency domain. First, video frames are split into patches and each patch is transformed into spectral maps in which each channel represents a frequency band. It permits a fine-grained self-attention on each frequency band, so that real visual texture can be distinguished from artifacts. Second, a novel dual frequency attention (DFA) mechanism is proposed to capture the global frequency relations and local frequency relations, which can handle different complicated degradation processes in real-world scenarios. Third, we explore different self-attention schemes for video processing in the frequency domain and discover that a ``divided attention'' which conducts a joint space-frequency attention before applying temporal-frequency attention, leads to the best video enhancement quality. Extensive experiments on three widely-used VSR datasets show that FTVSR outperforms state-of-the-art methods on different low-quality videos with clear visual margins. Code and pre-trained models are available at https://github.com/researchmm/FTVSR.

📄 PDF Abstract BibTeX arXiv:2212.14046

Code (1)

researchmm/ftvsr 공식 구현 pytorch

Tasks

Super-ResolutionVideo EnhancementVideo Super-Resolution

Similar Papers 제목 키워드 기반

Learning Spatiotemporal Frequency-Transformer for Compressed Video Super-Resolution

2022-08-05 · Zhongwei Qiu, Huan Yang, Jianlong Fu, Dongmei Fu

Compressed video super-resolution (VSR) aims to restore high-resolution frames from compressed low-resolution counterparts. Most recent VSR approaches often enhance an input frame by borrowing relevant textures from neig…

Super-ResolutionVideo EnhancementVideo Super-Resolution

Frequency-Aware Spatiotemporal Transformers for Video Inpainting Detection

2021-01-01 · ICCV 2021 10 · Bingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu 외

In this paper, we propose a frequency-aware spatiotemporal transformers for deep In this paper, we propose a Frequency-Aware Spatiotemporal Transformer (FAST) for video inpainting detection, which aims to simultaneou…

DecoderVideo Inpainting

DiTVR: Zero-Shot Diffusion Transformer for Video Restoration

2025-08-11 · Sicheng Gao, Nancy Mehta, Zongwei Wu, Radu Timofte arxiv

Video restoration aims to reconstruct high quality video sequences from low quality inputs, addressing tasks such as super resolution, denoising, and deblurring. Traditional regression based methods often produce unreali…

Video Restoration

MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion

2024-10-10 · Onkar Susladkar, Jishu Sen Gupta, Chirag Sehgal, Sparsh Mittal 외

The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video p…

Denoisingparameter-efficient fine-tuningQuantizationText-to-Video Generation+3

Self-supervised Video Transformer

2021-12-02 · CVPR 2022 1 · Kanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan 외

In this paper, we propose self-supervised training for video transformers using unlabeled video data. From a given video, we create local and global spatiotemporal views with varying spatial sizes and frame rates. Our se…

Action ClassificationAction RecognitionAction Recognition In VideosSelf-Supervised Action Recognition Linear