paper-with-me

Papers

RankDVQA-mini: Knowledge Distillation-Driven Deep Video Quality Assessment

2023-12-14 · Chen Feng, Duolikun Danier, Haoran Wang, Fan Zhang, Benoit Vallade, Alex Mackin, David Bull

Deep learning-based video quality assessment (deep VQA) has demonstrated significant potential in surpassing conventional metrics, with promising improvements in terms of correlation with human perception. However, the practical deployment of such deep VQA models is often limited due to their high computational complexity and large memory requirements. To address this issue, we aim to significantly reduce the model size and runtime of one of the state-of-the-art deep VQA methods, RankDVQA, by employing a two-phase workflow that integrates pruning-driven model compression with multi-level knowledge distillation. The resulting lightweight full reference quality metric, RankDVQA-mini, requires less than 10% of the model parameters compared to its full version (14% in terms of FLOPs), while still retaining a quality prediction performance that is superior to most existing deep VQA methods. The source code of the RankDVQA-mini has been released at https://chenfeng-bristol.github.io/RankDVQA-mini/ for public evaluation.

📄 PDF Abstract BibTeX arXiv:2312.08864

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationModel CompressionVideo Quality AssessmentVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

RankDVQA: Deep VQA based on Ranking-inspired Hybrid Training

2022-02-17 · Chen Feng, Duolikun Danier, Fan Zhang, David Bull

In recent years, deep learning techniques have shown significant potential for improving video quality assessment (VQA), achieving higher correlation with subjective opinions compared to conventional approaches. However,…

Video Quality AssessmentVisual Question Answering (VQA)

ST-MFNet Mini: Knowledge Distillation-Driven Frame Interpolation

2023-02-16 · Crispian Morris, Duolikun Danier, Fan Zhang, Nantheera Anantrasirichai 외

Currently, one of the major challenges in deep learning-based video frame interpolation (VFI) is the large model sizes and high computational complexity associated with many high performance VFI approaches. In this paper…

Knowledge DistillationNetwork PruningVideo Frame Interpolation

Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation

2026-06-23 · Tianyu Zhu, Yingping Liang, Hesong Li, Ying Fu arxiv

Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing models are typically trained on 2D image or video datasets with naive s…

Referring Video Object SegmentationZero-shot GeneralizationImage SegmentationVideo Segmentation

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation

2026-04-16 · Xiangyu Liu, Feng Gao, Xiaomei Zhang, Yong Zhang 외 arxiv

Existing audio-driven video digital human generation models rely on multi-step denoising, resulting in substantial computational overhead that severely limits their deployment in real-world settings. While one-step disti…

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

2026-04-26 · Chunyu Li, Jiaye Li, Ruiqiao Mei, Haoyuan Xia 외 arxiv

Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio-visual diffusion models remain too slow…