paper-with-me

홈 › Papers

Capturing Co-existing Distortions in User-Generated Content for No-reference Video Quality Assessment

2023-07-31 · Kun Yuan, Zishang Kong, Chuanchuan Zheng, Ming Sun, Xing Wen

Video Quality Assessment (VQA), which aims to predict the perceptual quality of a video, has attracted raising attention with the rapid development of streaming media technology, such as Facebook, TikTok, Kwai, and so on. Compared with other sequence-based visual tasks (\textit{e.g.,} action recognition), VQA faces two under-estimated challenges unresolved in User Generated Content (UGC) videos. \textit{First}, it is not rare that several frames containing serious distortions (\textit{e.g.,}blocking, blurriness), can determine the perceptual quality of the whole video, while other sequence-based tasks require more frames of equal importance for representations. \textit{Second}, the perceptual quality of a video exhibits a multi-distortion distribution, due to the differences in the duration and probability of occurrence for various distortions. In order to solve the above challenges, we propose \textit{Visual Quality Transformer (VQT)} to extract quality-related sparse features more efficiently. Methodologically, a Sparse Temporal Attention (STA) is proposed to sample keyframes by analyzing the temporal correlation between frames, which reduces the computational complexity from $O(T^2)$ to $O(T \log T)$. Structurally, a Multi-Pathway Temporal Network (MPTN) utilizes multiple STA modules with different degrees of sparsity in parallel, capturing co-existing distortions in a video. Experimentally, VQT demonstrates superior performance than many \textit{state-of-the-art} methods in three public no-reference VQA datasets. Furthermore, VQT shows better performance in four full-reference VQA datasets against widely-adopted industrial algorithms (\textit{i.e.,} VMAF and AVQT).

📄 PDF Abstract BibTeX arXiv:2307.16813

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionBlockingVideo Quality AssessmentVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Position-Wise Feed-Forward Layer 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

CHUG: Crowdsourced User-Generated HDR Video Quality Dataset

2025-10-10 · Shreshth Saini, Alan C. Bovik, Neil Birkbeck, Yilin Wang 외 arxiv

High Dynamic Range (HDR) videos enhance visual experiences with superior brightness, contrast, and color depth. The surge of User-Generated Content (UGC) on platforms like YouTube and TikTok introduces unique challenges …

Video Quality Assessment

Priorformer: A UGC-VQA Method with content and distortion priors

2024-06-24 · Yajing Pei, Shiyu Huang, Yiting Lu, Xin Li 외

User Generated Content (UGC) videos are susceptible to complicated and variant degradations and contents, which prevents the existing blind video quality assessment (BVQA) models from good performance since the lack of t…

Video Quality AssessmentVisual Question Answering (VQA)

Audio-Visual Quality Assessment for User Generated Content: Database and Method

2023-03-04 · Yuqin Cao, Xiongkuo Min, Wei Sun, XiaoPing Zhang 외

With the explosive increase of User Generated Content (UGC), UGC video quality assessment (VQA) becomes more and more important for improving users' Quality of Experience (QoE). However, most existing UGC VQA studies onl…

Video Quality AssessmentVisual Question Answering (VQA)

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

2025-01-30 · Yuqin Cao, Xiongkuo Min, Yixuan Gao, Wei Sun 외

Many video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual qualit…

TruthSR: Trustworthy Sequential Recommender Systems via User-generated Multimodal Content

2024-04-26 · Meng Yan, Haibin Huang, Ying Liu, Juan Zhao 외

Sequential recommender systems explore users' preferences and behavioral patterns from their historically generated data. Recently, researchers aim to improve sequential recommendation by utilizing massive user-generated…

Recommendation SystemsSequential Recommendation