paper-with-me

홈 › Papers

FVQ: A Large-Scale Dataset and A LMM-based Method for Face Video Quality Assessment

2025-04-12 · Sijing Wu, Yunhao Li, Ziwen Xu, Yixuan Gao, Huiyu Duan, Wei Sun, Guangtao Zhai

Face video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is particularly sensitive to human faces. However, FVQA is rarely explored due to the lack of large-scale FVQA datasets. To fill this gap, we present the first large-scale in-the-wild FVQA dataset, FVQ-20K, which contains 20,000 in-the-wild face videos together with corresponding mean opinion score (MOS) annotations. Along with the FVQ-20K dataset, we further propose a specialized FVQA method named FVQ-Rater to achieve human-like rating and scoring for face video, which is the first attempt to explore the potential of large multimodal models (LMMs) for the FVQA task. Concretely, we elaborately extract multi-dimensional features including spatial features, temporal features, and face-specific features (i.e., portrait features and face embeddings) to provide comprehensive visual information, and take advantage of the LoRA-based instruction tuning technique to achieve quality-specific fine-tuning, which shows superior performance on both FVQ-20K and CFVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the FVQ-20K dataset and FVQ-Rater method in promoting the development of FVQA.

📄 PDF Abstract BibTeX arXiv:2504.09255

Code (1)

wsj-sjtu/fvq 공식 구현

Tasks

Video Quality AssessmentVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

FaceVid-1K: A Large-Scale High-Quality Multiracial Human Face Video Dataset

2024-09-23 · Donglin Di, He Feng, Wenzhang Sun, Yongjia Ma 외

Generating talking face videos from various conditions has recently become a highly popular research area within generative tasks. However, building a high-quality face video generation model requires a well-performing p…

Image GenerationUnconditional Video GenerationVideo Generation

DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery Detection

2020-01-09 · CVPR 2020 6 · Liming Jiang, Ren Li, Wayne Wu, Chen Qian 외

We present our on-going effort of constructing a large-scale benchmark for face forgery detection. The first version of this benchmark, DeeperForensics-1.0, represents the largest face forgery detection dataset by far, w…

DiversityFace SwappingVideo Forensics

CelebV-Text: A Large-Scale Facial Text-Video Dataset

2023-03-26 · CVPR 2023 1 · Jianhui Yu, Hao Zhu, Liming Jiang, Chen Change Loy 외

Text-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and h…

Text GenerationText-to-Video GenerationVideo Generation

Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics

2019-09-27 · CVPR 2020 6 · Yuezun Li, Xin Yang, Pu Sun, Honggang Qi 외

AI-synthesized face-swapping videos, commonly known as DeepFakes, is an emerging problem threatening the trustworthiness of online information. The need to develop and evaluate DeepFake detection algorithms calls for lar…

DeepFake DetectionFace Swapping

Unsupervised Domain Adaptation for Face Recognition in Unlabeled Videos

2017-08-07 · ICCV 2017 10 · Kihyuk Sohn, Sifei Liu, Guangyu Zhong, Xiang Yu 외

Despite rapid advances in face recognition, there remains a clear gap between the performance of still image-based face recognition and video-based face recognition, due to the vast difference in visual quality between t…

Data AugmentationDomain AdaptationFace RecognitionUnsupervised Domain Adaptation