paper-with-me

홈 › Papers

Small Clips, Big Gains: Learning Long-Range Refocused Temporal Information for Video Super-Resolution

2025-05-04 · Xingyu Zhou, Wei Long, Jingbo Lu, Shiyin Jiang, Weiyi You, Haifeng Wu, Shuhang Gu

Video super-resolution (VSR) can achieve better performance compared to single image super-resolution by additionally leveraging temporal information. In particular, the recurrent-based VSR model exploits long-range temporal information during inference and achieves superior detail restoration. However, effectively learning these long-term dependencies within long videos remains a key challenge. To address this, we propose LRTI-VSR, a novel training framework for recurrent VSR that efficiently leverages Long-Range Refocused Temporal Information. Our framework includes a generic training strategy that utilizes temporal propagation features from long video clips while training on shorter video clips. Additionally, we introduce a refocused intra&inter-frame transformer block which allows the VSR model to selectively prioritize useful temporal information through its attention module while further improving inter-frame information utilization in the FFN module. We evaluate LRTI-VSR on both CNN and transformer-based VSR architectures, conducting extensive ablation studies to validate the contribution of each component. Experiments on long-video test sets demonstrate that LRTI-VSR achieves state-of-the-art performance while maintaining training and computational efficiency.

📄 PDF Abstract BibTeX arXiv:2505.02159

Code (1)

labshuhanggu/lrti-vsr 공식 구현

Tasks

Computational EfficiencyImage Super-ResolutionSuper-ResolutionVideo Super-Resolution

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Light Field Synthesis by Training Deep Network in the Refocused Image Domain

2019-10-14 · Chang-Le Liu, Kuang-Tsu Shih, Jiun-Woei Huang, Homer H. Chen

Light field imaging, which captures spatio-angular information of incident light on image sensor, enables many interesting applications like image refocusing and augmented reality. However, due to the limited sensor reso…

SSIM

Revisiting Kernel Temporal Segmentation as an Adaptive Tokenizer for Long-form Video Understanding

2023-09-20 · Mohamed Afham, Satya Narayan Shukla, Omid Poursaeed, Pengchuan Zhang 외

While most modern video understanding models operate on short-range clips, real-world videos are often several minutes long with semantically consistent segments of variable length. A common approach to process long vide…

Action LocalizationFormTemporal Action LocalizationVideo Classification+1

Model Extraction Attack against Self-supervised Speech Models

2022-11-29 · Tsu-Yuan Hsu, Chen-An Li, Tung-Yu Wu, Hung-Yi Lee

Self-supervised learning (SSL) speech models generate meaningful representations of given clips and achieve incredible performance across various downstream tasks. Model extraction attack (MEA) often refers to an adversa…

modelModel extractionSelf-Supervised Learning

Learning to See Through with Events

2022-12-05 · Lei Yu, Xiang Zhang, Wei Liao, Wen Yang 외

Although synthetic aperture imaging (SAI) can achieve the seeing-through effect by blurring out off-focus foreground occlusions while recovering in-focus occluded scenes from multi-view images, its performance is often d…

Removing Dynamic Objects for Static Scene Reconstruction using Light Fields

2020-03-24 · Pushyami Kaveti, Sammie Katt, Hanumant Singh

There is a general expectation that robots should operate in environments that consist of static and dynamic entities including people, furniture and automobiles. These dynamic environments pose challenges to visual simu…

GPUSemantic SegmentationSimultaneous Localization and Mapping