paper-with-me

홈 › Papers

Analysis of Video Quality Datasets via Design of Minimalistic Video Quality Models

2023-07-26 · Wei Sun, Wen Wen, Xiongkuo Min, Long Lan, Guangtao Zhai, Kede Ma

Blind video quality assessment (BVQA) plays an indispensable role in monitoring and improving the end-users' viewing experience in various real-world video-enabled media applications. As an experimental field, the improvements of BVQA models have been measured primarily on a few human-rated VQA datasets. Thus, it is crucial to gain a better understanding of existing VQA datasets in order to properly evaluate the current progress in BVQA. Towards this goal, we conduct a first-of-its-kind computational analysis of VQA datasets via designing minimalistic BVQA models. By minimalistic, we restrict our family of BVQA models to build only upon basic blocks: a video preprocessor (for aggressive spatiotemporal downsampling), a spatial quality analyzer, an optional temporal quality analyzer, and a quality regressor, all with the simplest possible instantiations. By comparing the quality prediction performance of different model variants on eight VQA datasets with realistic distortions, we find that nearly all datasets suffer from the easy dataset problem of varying severity, some of which even admit blind image quality assessment (BIQA) solutions. We additionally justify our claims by contrasting our model generalizability on these VQA datasets, and by ablating a dizzying set of BVQA design choices related to the basic building blocks. Our results cast doubt on the current progress in BVQA, and meanwhile shed light on good practices of constructing next-generation VQA datasets and models.

📄 PDF Abstract BibTeX arXiv:2307.13981

Code (1)

sunwei925/minimalisticvqa 공식 구현 pytorch

Tasks

Image Quality AssessmentNo-Reference Image Quality AssessmentVideo Quality AssessmentVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Minimalistic Video Saliency Prediction via Efficient Decoder & Spatio Temporal Action Cues

2025-02-01 · Rohit Girmaji, Siddharth Jain, Bhav Beri, Sarthak Bansal 외

This paper introduces ViNet-S, a 36MB model based on the ViNet architecture with a U-Net design, featuring a lightweight decoder that significantly reduces model size and parameters without compromising performance. Addi…

Action ClassificationAction LocalizationDecoderSaliency Prediction+3

Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery

2025-09-02 · Xinrui Gong, Oliver Hahn, Christoph Reich, Krishnakant Singh 외 arxiv

Unsupervised multi-object discovery (MOD) aims to detect and localize distinct object instances in visual scenes without any form of human supervision. Recent approaches leverage object-centric learning (OCL) and motion …

Multi-object discoveryMotion Segmentation

DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection

2025-06-06 · Marcel Klemt, Carlotta Segna, Anna Rohrbach

Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to…

BenchmarkingDeepFake DetectionFace SwappingMisinformation

Temporal ordering of clinical events

2015-04-14 · Azad Dehghan

This report describes a minimalistic set of methods engineered to anchor clinical events onto a temporal space. Specifically, we describe methods to extract clinical events (e.g., Problems, Treatments and Tests), tempora…

UTFPR at WMT 2018: Minimalistic Supervised Corpora Filtering for Machine Translation

2018-10-01 · WS 2018 10 · Gustavo Paetzold

We present the UTFPR systems at the WMT 2018 parallel corpus filtering task. Our supervised approach discerns between good and bad translations by training classic binary classification models over an artificially produc…

Binary ClassificationClassificationGeneral ClassificationLanguage Modeling+4