paper-with-me

Papers

Scale-Adaptive Feature Aggregation for Efficient Space-Time Video Super-Resolution

2023-10-26 · Zhewei Huang, Ailin Huang, Xiaotao Hu, Chen Hu, Jun Xu, Shuchang Zhou

The Space-Time Video Super-Resolution (STVSR) task aims to enhance the visual quality of videos, by simultaneously performing video frame interpolation (VFI) and video super-resolution (VSR). However, facing the challenge of the additional temporal dimension and scale inconsistency, most existing STVSR methods are complex and inflexible in dynamically modeling different motion amplitudes. In this work, we find that choosing an appropriate processing scale achieves remarkable benefits in flow-based feature propagation. We propose a novel Scale-Adaptive Feature Aggregation (SAFA) network that adaptively selects sub-networks with different processing scales for individual samples. Experiments on four public STVSR benchmarks demonstrate that SAFA achieves state-of-the-art performance. Our SAFA network outperforms recent state-of-the-art methods such as TMNet and VideoINR by an average improvement of over 0.5dB on PSNR, while requiring less than half the number of parameters and only 1/3 computational costs.

📄 PDF Abstract BibTeX arXiv:2310.17294

Code (1)

megvii-research/wacv2024-safa 공식 구현 pytorch

Tasks

Space-time Video Super-resolutionSuper-ResolutionVideo Frame InterpolationVideo Super-Resolution

Similar Papers 제목 키워드 기반

Content-Aware Inter-Scale Cost Aggregation for Stereo Matching

2020-06-05 · Chengtang Yao, Yunde Jia, Huijun Di, Yuwei Wu 외

Cost aggregation is a key component of stereo matching for high-quality depth estimation. Most methods use multi-scale processing to downsample cost volume for proper context information, but will cause loss of details w…

Depth EstimationStereo Matching

TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object Detection

2026-08-27 · Su Wang, Yaochen Li, Min Yang, Jiaohao Nie 외 arxiv

Most single-stage 3D object detectors complete different tasks with the same extracted features. Nevertheless, it is impossible to project features into a common space that is adaptive for all the tasks. We present a nov…

3D Object Detection

Learnable Query Aggregation with KV Routing for Cross-view Geo-localisation

2025-12-30 · Hualin Ye, Bingxi Liu, Jixiang Du, Yu Qin 외 arxiv

Cross-view geo-localisation (CVGL) aims to estimate the geographic location of a query image by matching it with images from a large-scale database. However, the significant view-point discrepancies present considerable …

Cross-View Geo-Localisation

FANet: Quality-Aware Feature Aggregation Network for Robust RGB-T Tracking

2018-11-24 · Yabin Zhu, Chenglong Li, Bin Luo, Jin Tang

This paper investigates how to perform robust visual tracking in adverse and challenging conditions using complementary visual and thermal infrared data (RGBT tracking). We propose a novel deep network architecture calle…

Rgb-T TrackingVisual Tracking

CMSA-Net: Causal Multi-scale Aggregation with Adaptive Multi-source Reference for Video Polyp Segmentation

2026-02-26 · Tong Wang, Yaolei Qi, Siwen Wang, Imran Razzak 외 arxiv

Video polyp segmentation (VPS) is an important task in computer-aided colonoscopy, as it helps doctors accurately locate and track polyps during examinations. However, VPS remains challenging because polyps often look si…

Video Polyp Segmentation