Learning Generalized Spatial-Temporal Deep Feature Representation for No-Reference Video Quality Assessment
In this work, we propose a no-reference video quality assessment method, aiming to achieve high-generalization capability in cross-content, -resolution and -frame rate quality prediction. In particular, we evaluate the quality of a video by learning effective feature representations in spatial-temporal domain. In the spatial domain, to tackle the resolution and content variations, we impose the Gaussian distribution constraints on the quality features. The unified distribution can significantly reduce the domain gap between different video samples, resulting in a more generalized quality feature representation. Along the temporal dimension, inspired by the mechanism of visual perception, we propose a pyramid temporal aggregation module by involving the short-term and long-term memory to aggregate the frame-level quality. Experiments show that our method outperforms the state-of-the-art methods on cross-dataset settings, and achieves comparable performance on intra-dataset configurations, demonstrating the high-generalization capability of the proposed method.
Code (1)
Tasks
Video Quality AssessmentSimilar Papers 제목 키워드 기반
Pyramid Spatial-Temporal Aggregation for Video-Based Person Re-Identification
Video-based person re-identification aims to associate the video clips of the same person across multiple non-overlapping cameras. Spatial-temporal representations can provide richer and complementary information bet…
Person Re-IdentificationVideo-Based Person Re-IdentificationVideo Quality Assessment Based on Swin TransformerV2 and Coarse to Fine Strategy
The objective of non-reference video quality assessment is to evaluate the quality of distorted video without access to reference high-definition references. In this study, we introduce an enhanced spatial perception mod…
Image Quality AssessmentVideo Quality AssessmentVisual Question Answering (VQA)ST-GREED: Space-Time Generalized Entropic Differences for Frame Rate Dependent Video Quality Prediction
We consider the problem of conducting frame rate dependent video quality assessment (VQA) on videos of diverse frame rates, including high frame rate (HFR) videos. More generally, we study how perceptual quality is affec…
Video Quality AssessmentVisual Question Answering (VQA)Blockwise Temporal-Spatial Pathway Network
Algorithms for video action recognition should consider not only spatial information but also temporal relations, which remains challenging. We propose a 3D-CNN-based action recognition model, called the blockwise tempor…
Action RecognitionTemporal Action LocalizationTSception: Capturing Temporal Dynamics and Spatial Asymmetry from EEG for Emotion Recognition
The high temporal resolution and the asymmetric spatial activations are essential attributes of electroencephalogram (EEG) underlying emotional processes in the brain. To learn the temporal dynamics and spatial asymmetry…
EEGEmotion Recognition