paper-with-me

홈 › Papers

VariFace: Fair and Diverse Synthetic Dataset Generation for Face Recognition

2024-12-09 · Michael Yeung, Toya Teramoto, Songtao Wu, Tatsuo Fujiwara, Kenji Suzuki, Tamaki Kojima

The use of large-scale, web-scraped datasets to train face recognition models has raised significant privacy and bias concerns. Synthetic methods mitigate these concerns and provide scalable and controllable face generation to enable fair and accurate face recognition. However, existing synthetic datasets display limited intraclass and interclass diversity and do not match the face recognition performance obtained using real datasets. Here, we propose VariFace, a two-stage diffusion-based pipeline to create fair and diverse synthetic face datasets to train face recognition models. Specifically, we introduce three methods: Face Recognition Consistency to refine demographic labels, Face Vendi Score Guidance to improve interclass diversity, and Divergence Score Conditioning to balance the identity preservation-intraclass diversity trade-off. When constrained to the same dataset size, VariFace considerably outperforms previous synthetic datasets (0.9200 $\rightarrow$ 0.9405) and achieves comparable performance to face recognition models trained with real data (Real Gap = -0.0065). In an unconstrained setting, VariFace not only consistently achieves better performance compared to previous synthetic methods across dataset sizes but also, for the first time, outperforms the real dataset (CASIA-WebFace) across six evaluation datasets. This sets a new state-of-the-art performance with an average face verification accuracy of 0.9567 (Real Gap = +0.0097) across LFW, CFP-FP, CPLFW, AgeDB, and CALFW datasets and 0.9366 (Real Gap = +0.0380) on the RFW dataset.

📄 PDF Abstract BibTeX arXiv:2412.06235

Code (0)

등록된 구현이 없습니다.

Tasks

Dataset GenerationDiversityFace GenerationFace RecognitionFace Verification

Similar Papers 제목 키워드 기반

FairDD: Fair Dataset Distillation via Synchronized Matching

2024-11-29 · Qihang Zhou, Shenhao Fang, Shibo He, Wenchao Meng 외

Condensing large datasets into smaller synthetic counterparts has demonstrated its promise for image classification. However, previous research has overlooked a crucial concern in image recognition: ensuring that models …

Dataset DistillationFairnessimage-classificationImage Classification

Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise

2025-10-20 · Paweł Borsukiewicz, Fadi Boutros, Iyiola E. Olatunji, Charles Beumier 외 arxiv

The deployment of facial recognition systems has created an ethical dilemma: achieving high accuracy requires massive datasets of real faces collected without consent, leading to dataset retractions and potential legal l…

FairFinGAN: Fairness-aware Synthetic Financial Data Generation

2026-03-05 · Tai Le Quy, Dung Nguyen Tuan, Trung Nguyen Thanh, Duy Tran Cong 외 arxiv

Financial datasets often suffer from bias that can lead to unfair decision-making in automated systems. In this work, we propose FairFinGAN, a WGAN-based framework designed to generate synthetic financial data while miti…

Can Synthetic Data be Fair and Private? A Comparative Study of Synthetic Data Generation and Fairness Algorithms

2025-01-03 · Qinyi Liu, Oscar Deho, Farhad Vadiee, Mohammad Khalil 외

The increasing use of machine learning in learning analytics (LA) has raised significant concerns around algorithmic fairness and privacy. Synthetic data has emerged as a dual-purpose tool, enhancing privacy and improvin…

FairnessSynthetic Data Generation

Data-Driven Fairness Generalization for Deepfake Detection

2024-12-21 · Uzoamaka Ezeakunne, Chrisantus Eze, Xiuwen Liu

Despite the progress made in deepfake detection research, recent studies have shown that biases in the training data for these detectors can result in varying levels of performance across different demographic groups, su…

DeepFake DetectionFace SwappingFairnessImage Manipulation+3