paper-with-me

홈 › Papers

FDIM: A Feature-distance-based Generic Video Quality Metric for Versatile Codecs

2026-04-27 · Jiayi Wang, Lichun Zhang, Xiaoqi Zhuang, Jiaqi Zhang, Lu Yu, Yin Zhao arxiv

Video technology is advancing toward Ultra High Definition (UHD) and High Dynamic Range (HDR), which intensifies the need for higher compression efficiency for these high-specification videos. Beyond advances in traditional codecs, neural video codecs (NVCs) have attracted significant research attention and have evolved rapidly over the past few years. The coding artifacts of NVCs often exhibit content-varying and generative characteristics, which differ from those of conventional codecs and are challenging for traditional video quality assessment (VQA) methods to capture. Therefore, VQA metrics are required to generalize across different codecs, content types, and dynamic ranges to better support video codec research and evaluation. In this paper, we propose FDIM, a feature-distance-based generic video quality metric for both traditional and neural video codecs across SDR and HDR formats. FDIM employs a hybrid architecture that integrates deep and hand-crafted features. The deep feature component learns multi-scale representations to capture distortions ranging from structural and textural fidelity degradation to high-level semantic deviations, while the hand-crafted feature component provides stable complementary cues to improve overall generalization. We trained FDIM on a large-scale subjective quality assessment dataset (DCVQA) consisting of over 16k video sequences encoded by traditional block-based hybrid video codecs and end-to-end perceptually optimized neural video codecs. Extensive experiments on ten SDR/HDR VQA datasets containing diverse, previously unseen codecs demonstrate that FDIM achieves strong generalization and high correlation with subjective assessment. The source code for FDIM and the DCVQA validation set will be released at https://github.com/MCL-ZJU/FDIM.

📄 PDF Abstract BibTeX arXiv:2604.24123

Code (0)

등록된 구현이 없습니다.

Tasks

Video Quality Assessment

Similar Papers 제목 키워드 기반

Geodesic Distance Histogram Feature for Video Segmentation

2017-03-31 · Hieu Le, Vu Nguyen, Chen-Ping Yu, Dimitris Samaras

This paper proposes a geodesic-distance-based feature that encodes global information for improved video segmentation algorithms. The feature is a joint histogram of intensity and geodesic distances, where the geodesic d…

SegmentationSuperpixelsVideo SegmentationVideo Semantic Segmentation

Fréchet Video Motion Distance: A Metric for Evaluating Motion Consistency in Videos

2024-07-23 · Jiahe Liu, Youran Qu, Qi Yan, Xiaohui Zeng 외

Significant advancements have been made in video generative models recently. Unlike image generation, video generation presents greater challenges, requiring not only generating high-quality frames but also ensuring temp…

Image GenerationPoint TrackingVideo GenerationVideo Quality Assessment+1

IAE-VTG: Interaction-Aligned Action-Entity Video Temporal Grounding

2026-09-09 · Shiwen Zhao, Qi Zhang, Sezer Karaoglu, Theo Gevers 외 arxiv

Video Temporal Grounding (VTG) localizes the video segment that matches a natural-language query. Many queries describe an action performed by a particular entity. Existing methods often encode the query as a whole or us…

LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning

2023-06-17 · Yunlong Tang, Jinrui Zhang, Xiangchen Wang, Teng Wang 외

Our winning entry for the CVPR 2023 Generic Event Boundary Captioning (GEBC) competition is detailed in this paper. Unlike conventional video captioning tasks, GEBC demands that the captioning model possess an understand…

Boundary CaptioningLanguage ModelingLanguage ModellingLarge Language Model+1

MARLIN: Masked Autoencoder for facial video Representation LearnINg

2022-11-12 · CVPR 2023 1 · Zhixi Cai, Shreya Ghosh, Kalin Stefanov, Abhinav Dhall 외

This paper proposes a self-supervised approach to learn universal facial representations from videos, that can transfer across a variety of facial analysis tasks such as Facial Attribute Recognition (FAR), Facial Express…

Action ClassificationAttributeDeepFake DetectionEmotion Classification+8