paper-with-me

홈 › Papers

A Camera-Native Talking-Head Video Dataset for Various Computer Vision Tasks

2026-03-23 · Babak Naderi, Ross Cutler, Nabakumar Singh Khongbantabam arxiv

Talking-head videos constitute a predominant content type in real-time communication, yet publicly available datasets for video processing research in this domain remain scarce and limited in signal fidelity. In this paper, we open-source a camera-native dataset of 838 talking-head recordings (approximately 210 minutes), each 15s in duration, captured from 799 participants across 443 camera-code categories in their natural environments. All recordings are stored using the FFV1 lossless codec, preserving the camera-native signal---uncompressed (24.7%) or MJPEG-encoded (75.3%)---without additional lossy processing. Each recording is annotated with a Mean Opinion Score (MOS) and ten perceptual quality tokens that jointly explain 64.4% of the MOS variance. From this corpus, we curate a stratified benchmarking subset of 120 clips in three content conditions: original, background blur, and background replacement. Codec efficiency evaluation across four datasets and four codecs, namely H.264, H.265, H.266, and AV1, yields VMAF BD-rate savings up to $-71.3%$ (H.266) relative to H.264, with significant encoder$\times$dataset ($η_p^2 = .112$) and encoder$\times$content condition ($η_p^2 = .149$) interactions, demonstrating that both content type and background processing affect compression efficiency. A preliminary super-resolution evaluation with four SR models confirms that the dataset significantly affects absolute performance while preserving model rankings, demonstrating applicability beyond codec benchmarking. The dataset offers 5$\times$ the scale of the largest prior talking-head webcam dataset (838 vs. 160 clips) and preserves the camera-native signal without additional lossy compression, establishing a resource for benchmarking video compression, super-resolution, quality assessment, and enhancement models in real-time communication.

📄 PDF Abstract BibTeX arXiv:2603.26763

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Toward Fine-Grained Facial Control in 3D Talking Head Generation

2026-02-10 · Shaoyang Xie, Xiaofeng Cong, Baosheng Yu, Zhipeng Gui 외 arxiv

Audio-driven talking head generation is a core component of digital avatars, and 3D Gaussian Splatting has shown strong performance in real-time rendering of high-fidelity talking heads. However, achieving precise contro…

Talking Head Generation

One-Shot Pose-Driving Face Animation Platform

2024-07-12 · He Feng, Donglin Di, Yongjia Ma, Wei Chen 외

The objective of face animation is to generate dynamic and expressive talking head videos from a single reference face, utilizing driving conditions derived from either video or audio inputs. Current approaches often req…

TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

2023-04-01 · Yifeng Ma, Suzhen Wang, Yu Ding, Bowen Ma 외

Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which m…

2D Semantic Segmentation task 3 (25 classes)Talking Head Generation

Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose

2020-02-24 · Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao 외

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this proble…

3D Face AnimationVideo Generation

AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person

2021-08-09 · Xinsheng Wang, Qicong Xie, Jihua Zhu, Lei Xie 외

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic me…

Talking Head Generationtext-to-speechText to Speech