paper-with-me

Papers

Representative & Fair Synthetic Data

2021-04-07 · Paul Tiwald, Alexandra Ebert, Daniel T. Soukup

Algorithms learn rules and associations based on the training data that they are exposed to. Yet, the very same data that teaches machines to understand and predict the world, contains societal and historic biases, resulting in biased algorithms with the risk of further amplifying these once put into use for decision support. Synthetic data, on the other hand, emerges with the promise to provide an unlimited amount of representative, realistic training samples, that can be shared further without disclosing the privacy of individual subjects. We present a framework to incorporate fairness constraints into the self-supervised learning process, that allows to then simulate an unlimited amount of representative as well as fair synthetic data. This framework provides a handle to govern and control for privacy as well as for bias within AI at its very source: the training data. We demonstrate the proposed approach by amending an existing generative model architecture and generating a representative as well as fair version of the UCI Adult census data set. While the relationships between attributes are faithfully retained, the gender and racial biases inherent in the original data are controlled for. This is further validated by comparing propensity scores of downstream predictive models that are trained on the original data versus the fair synthetic data. We consider representative & fair synthetic data a promising future building block to teach algorithms not on historic worlds, but rather on the worlds that we strive to live in.

📄 PDF Abstract BibTeX arXiv:2104.03007

Code (0)

등록된 구현이 없습니다.

Tasks

FairnessSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Fair Wasserstein Coresets

2023-11-09 · Zikai Xiong, Niccolò Dalmasso, Shubham Sharma, Freddy Lecue 외

Data distillation and coresets have emerged as popular approaches to generate a smaller representative set of samples for downstream learning tasks to handle large-scale datasets. At the same time, machine learning is be…

ClusteringDecision MakingFairness

MedEqualizer: A Framework Investigating Bias in Synthetic Medical Data and Mitigation via Augmentation

2025-11-02 · Sama Salarian, Yue Zhang, Swati Padhee, Srinivasan Parthasarathy arxiv

Synthetic healthcare data generation presents a viable approach to enhance data accessibility and support research by overcoming limitations associated with real-world medical datasets. However, ensuring fairness across …

Synthetic Data Generation

Data-Driven Fairness Generalization for Deepfake Detection

2024-12-21 · Uzoamaka Ezeakunne, Chrisantus Eze, Xiuwen Liu

Despite the progress made in deepfake detection research, recent studies have shown that biases in the training data for these detectors can result in varying levels of performance across different demographic groups, su…

DeepFake DetectionFace SwappingFairnessImage Manipulation+3

Assessment of Differentially Private Synthetic Data for Utility and Fairness in End-to-End Machine Learning Pipelines for Tabular Data

2023-10-30 · Mayana Pereira, Meghana Kshirsagar, Sumit Mukherjee, Rahul Dodhia 외

Differentially private (DP) synthetic data sets are a solution for sharing data while preserving the privacy of individual data providers. Understanding the effects of utilizing DP synthetic data in end-to-end machine le…

FairnessHumanitarianSynthetic Data Generation

Beyond Internal Data: Constructing Complete Datasets for Fairness Testing

2025-07-24 · Varsha Ramineni, Hossein A. Rahmani, Emine Yilmaz, David Barber arxiv

As AI becomes prevalent in high-risk domains and decision-making, it is essential to test for potential harms and biases. This urgency is reflected by the global emergence of AI regulations that emphasise fairness and ad…