paper-with-me

홈 › Papers

Fairness-Optimized Synthetic EHR Generation for Arbitrary Downstream Predictive Tasks

2024-06-04 · Mirza Farhan Bin Tarek, Raphael Poulain, Rahmatollah Beheshti

Among various aspects of ensuring the responsible design of AI tools for healthcare applications, addressing fairness concerns has been a key focus area. Specifically, given the wide spread of electronic health record (EHR) data and their huge potential to inform a wide range of clinical decision support tasks, improving fairness in this category of health AI tools is of key importance. While such a broad problem (mitigating fairness in EHR-based AI models) has been tackled using various methods, task- and model-agnostic methods are noticeably rare. In this study, we aimed to target this gap by presenting a new pipeline that generates synthetic EHR data, which is not only consistent with (faithful to) the real EHR data but also can reduce the fairness concerns (defined by the end-user) in the downstream tasks, when combined with the real data. We demonstrate the effectiveness of our proposed pipeline across various downstream tasks and two different EHR datasets. Our proposed pipeline can add a widely applicable and complementary tool to the existing toolbox of methods to address fairness in health AI applications, such as those modifying the design of a downstream model. The codebase for our project is available at https://github.com/healthylaife/FairSynth

📄 PDF Abstract BibTeX arXiv:2406.02510

Code (1)

healthylaife/fairsynth 공식 구현

Tasks

Fairness

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

FairFinGAN: Fairness-aware Synthetic Financial Data Generation

2026-03-05 · Tai Le Quy, Dung Nguyen Tuan, Trung Nguyen Thanh, Duy Tran Cong 외 arxiv

Financial datasets often suffer from bias that can lead to unfair decision-making in automated systems. In this work, we propose FairFinGAN, a WGAN-based framework designed to generate synthetic financial data while miti…

Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation

2025-08-25 · Rishikesh Devanathan, Varun Nathan, Ayush Kumar arxiv

Synthetic data is increasingly critical for contact centers, where privacy constraints and data scarcity limit the availability of real conversations. However, generating synthetic dialogues that are realistic and useful…

Dialogue Generation

CuTS: Customizable Tabular Synthetic Data Generation

2023-07-07 · Mark Vero, Mislav Balunović, Martin Vechev

Privacy, data quality, and data sharing concerns pose a key limitation for tabular data applications. While generating synthetic data resembling the original distribution addresses some of these issues, most applications…

FairnessSynthetic Data GenerationTabular Data Generation

Assessment of Differentially Private Synthetic Data for Utility and Fairness in End-to-End Machine Learning Pipelines for Tabular Data

2023-10-30 · Mayana Pereira, Meghana Kshirsagar, Sumit Mukherjee, Rahul Dodhia 외

Differentially private (DP) synthetic data sets are a solution for sharing data while preserving the privacy of individual data providers. Understanding the effects of utilizing DP synthetic data in end-to-end machine le…

FairnessHumanitarianSynthetic Data Generation

Improving Performance, Robustness, and Fairness of Radiographic AI Models with Finely-Controllable Synthetic Data

2025-08-22 · Stefania L. Moroianu, Christian Bluethgen, Pierre Chambon, Mehdi Cherti 외 arxiv

Achieving robust performance and fairness across diverse patient populations remains a challenge in developing clinically deployable deep learning models for diagnostic imaging. Synthetic data generation has emerged as a…

Synthetic Data Generation