paper-with-me

홈 › Papers

Priorformer: A UGC-VQA Method with content and distortion priors

2024-06-24 · Yajing Pei, Shiyu Huang, Yiting Lu, Xin Li, Zhibo Chen

User Generated Content (UGC) videos are susceptible to complicated and variant degradations and contents, which prevents the existing blind video quality assessment (BVQA) models from good performance since the lack of the adapability of distortions and contents. To mitigate this, we propose a novel prior-augmented perceptual vision transformer (PriorFormer) for the BVQA of UGC, which boots its adaptability and representation capability for divergent contents and distortions. Concretely, we introduce two powerful priors, i.e., the content and distortion priors, by extracting the content and distortion embeddings from two pre-trained feature extractors. Then we adopt these two powerful embeddings as the adaptive prior tokens, which are transferred to the vision transformer backbone jointly with implicit quality features. Based on the above strategy, the proposed PriorFormer achieves state-of-the-art performance on three public UGC VQA datasets including KoNViD-1K, LIVE-VQC and YouTube-UGC.

📄 PDF Abstract BibTeX arXiv:2406.16297

Code (0)

등록된 구현이 없습니다.

Tasks

Video Quality AssessmentVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Multi-Head Attention 설명 없음
Residual Connection 설명 없음
Vision Transformer The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over…

Similar Papers 제목 키워드 기반

PriorFormer: A Transformer for Real-time Monocular 3D Human Pose Estimation with Versatile Geometric Priors

2025-08-21 · Mohamed Adjel, Vincent Bonnet arxiv

This paper proposes a new lightweight Transformer-based lifter that maps short sequences of human 2D joint positions to 3D poses using a single camera. The proposed model takes as input geometric priors including segment…

Monocular 3D Human Pose Estimation

HomoGen: Enhanced Video Inpainting via Homography Propagation and Diffusion

2025-01-01 · CVPR 2025 1 · Ding Ding, Yueming Pan, Ruoyu Feng, Qi Dai 외

In this paper, we present HomoGen, an enhanced video inpainting method based on homography propagation and diffusion models. HomoGen leverages homography registration to propagate contextual pixels as priors for gene…

DenoisingVideo Inpainting

Compressing Scene Dynamics: A Generative Approach

2024-10-13 · Shanzhi Yin, Zihan Zhang, Bolin Chen, Shiqi Wang 외

This paper proposes to learn generative priors from the motion patterns instead of video contents for generative video compression. The priors are derived from small motion dynamics in common scenes such as swinging tree…

DecoderVideo Compression

Atmospheric Turbulence Removal with Video Sequence Deep Visual Priors

2024-02-29 · P. Hill, N. Anantrasirichai, A. Achim, D. R. Bull

Atmospheric turbulence poses a challenge for the interpretation and visual perception of visual imagery due to its distortion effects. Model-based approaches have been used to address this, but such methods often suffer …

Self-Supervised Learning

Enhancing Video Inpainting with Aligned Frame Interval Guidance

2025-10-24 · Ming Xie, Junqiu Yu, Qiaole Dong, Xiangyang Xue 외 arxiv

Recent image-to-video (I2V) based video inpainting methods have made significant strides by leveraging single-image priors and modeling temporal consistency across masked frames. Nevertheless, these methods suffer from s…

Image InpaintingVideo Inpainting