paper-with-me

홈 › Papers

FaceVid-1K: A Large-Scale High-Quality Multiracial Human Face Video Dataset

2024-09-23 · Donglin Di, He Feng, Wenzhang Sun, Yongjia Ma, Hao Li, Wei Chen, Xiaofei Gou, Tonghua Su, Xun Yang

Generating talking face videos from various conditions has recently become a highly popular research area within generative tasks. However, building a high-quality face video generation model requires a well-performing pre-trained backbone, a key obstacle that universal models fail to adequately address. Most existing works rely on universal video or image generation models and optimize control mechanisms, but they neglect the evident upper bound in video quality due to the limited capabilities of the backbones, which is a result of the lack of high-quality human face video datasets. In this work, we investigate the unsatisfactory results from related studies, gather and trim existing public talking face video datasets, and additionally collect and annotate a large-scale dataset, resulting in a comprehensive, high-quality multiracial face collection named \textbf{FaceVid-1K}. Using this dataset, we craft several effective pre-trained backbone models for face video generation. Specifically, we conduct experiments with several well-established video generation models, including text-to-video, image-to-video, and unconditional video generation, under various settings. We obtain the corresponding performance benchmarks and compared them with those trained on public datasets to demonstrate the superiority of our dataset. These experiments also allow us to investigate empirical strategies for crafting domain-specific video generation tasks with cost-effective settings. We will make our curated dataset, along with the pre-trained talking face video generation models, publicly available as a resource contribution to hopefully advance the research field.

📄 PDF Abstract BibTeX arXiv:2410.07151

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationUnconditional Video GenerationVideo Generation

Similar Papers 제목 키워드 기반

Evidence for Hypodescent in Visual Semantic AI

2022-05-22 · Robert Wolfe, Mahzarin R. Banaji, Aylin Caliskan

We examine the state-of-the-art multimodal "visual semantic" model CLIP ("Contrastive Language Image Pretraining") for the rule of hypodescent, or one-drop rule, whereby multiracial people are more likely to be assigned …

MORPH

Mind the gap: how multiracial individuals get left behind when we talk about race, ethnicity, and ancestry in genomic research

2022-04-29 · Daphne O. Martschenko, Hannah Wand, Jennifer L. Young, Genevieve L. Wojcik

It is widely acknowledged that there is a diversity problem in genomics stemming from the vast underrepresentation of non-European genetic ancestry populations. While many challenges exist to address this gap, a major co…

Diversity

Real-time Facial Expression Recognition "In The Wild'' by Disentangling 3D Expression from Identity

2020-05-12 · Mohammad Rami Koujan, Luma Alharbawee, Giorgos Giannakakis, Nicolas Pugeault 외

Human emotions analysis has been the focus of many studies, especially in the field of Affective Computing, and is important for many applications, e.g. human-computer intelligent interaction, stress analysis, interactiv…

Emotion RecognitionFacial Expression RecognitionFacial Expression Recognition (FER)

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

2026-08-07 · Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou 외 hf

The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmark…

Face Swapping

Silicon Minds versus Human Hearts: The Wisdom of Crowds Beats the Wisdom of AI in Emotion Recognition

2025-08-12 · Mustafa Akben, Vinayaka Gude, Haya Ajjan arxiv

The ability to discern subtle emotional cues is fundamental to human social intelligence. As artificial intelligence (AI) becomes increasingly common, AI's ability to recognize and respond to human emotions is crucial fo…

Emotional IntelligenceEmotion Recognition