paper-with-me

홈 › Papers

Zoom-VQA: Patches, Frames and Clips Integration for Video Quality Assessment

2023-04-13 · Kai Zhao, Kun Yuan, Ming Sun, Xing Wen

Video quality assessment (VQA) aims to simulate the human perception of video quality, which is influenced by factors ranging from low-level color and texture details to high-level semantic content. To effectively model these complicated quality-related factors, in this paper, we decompose video into three levels (\ie, patch level, frame level, and clip level), and propose a novel Zoom-VQA architecture to perceive spatio-temporal features at different levels. It integrates three components: patch attention module, frame pyramid alignment, and clip ensemble strategy, respectively for capturing region-of-interest in the spatial dimension, multi-level information at different feature levels, and distortions distributed over the temporal dimension. Owing to the comprehensive design, Zoom-VQA obtains state-of-the-art results on four VQA benchmarks and achieves 2nd place in the NTIRE 2023 VQA challenge. Notably, Zoom-VQA has outperformed the previous best results on two subsets of LSVQ, achieving 0.8860 (+1.0%) and 0.7985 (+1.9%) of SRCC on the respective subsets. Adequate ablation studies further verify the effectiveness of each component. Codes and models are released in https://github.com/k-zha14/Zoom-VQA.

📄 PDF Abstract BibTeX arXiv:2304.06440

Code (1)

k-zha14/zoom-vqa 공식 구현 pytorch

Tasks

Video Quality AssessmentVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Real-World Video for Zoom Enhancement based on Spatio-Temporal Coupling

2023-06-24 · Zhiling Guo, Yinqiang Zheng, Haoran Zhang, Xiaodan Shi 외

In recent years, single-frame image super-resolution (SR) has become more realistic by considering the zooming effect and using real-world short- and long-focus image pairs. In this paper, we further investigate the feas…

Image Super-ResolutionSuper-Resolution

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction

2024-11-22 · CVPR 2025 1 · Huiwon Jang, Sihyun Yu, Jinwoo Shin, Pieter Abbeel 외

Efficient tokenization of videos remains a challenge in training vision models that can process long videos. One promising direction is to develop a tokenizer that can encode long video clips, as it would enable the toke…

Deep Learning for Saliency Prediction in Natural Video

2016-04-27 · Souad Chaabouni, Jenny Benois-Pineau, Ofer Hadar, Chokri Ben Amar

The purpose of this paper is the detection of salient areas in natural video by using the new deep learning techniques. Salient patches in video frames are predicted first. Then the predicted visual fixation maps are bui…

Deep LearningPredictionSaliency PredictionSpecificity

EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 2022: Team HNU-FPV Technical Report

2022-07-07 · Nie Lin, Minjie Cai

In this report, we present the technical details of our submission to the 2022 EPIC-Kitchens Unsupervised Domain Adaptation (UDA) Challenge. Existing UDA methods align the global features extracted from the whole video c…

Action RecognitionDomain AdaptationUnsupervised Domain AdaptationVideo Recognition

Zoom in to the details of human-centric videos

2020-05-27 · Guanghan Li, Yaping Zhao, Mengqi Ji, Xiaoyun Yuan 외

Presenting high-resolution (HR) human appearance is always critical for the human-centric videos. However, current imagery equipment can hardly capture HR details all the time. Existing super-resolution algorithms barely…

Pose EstimationSuper-Resolution