paper-with-me

Papers

You Don't Have to Be Perfect to Be Amazing: Unveil the Utility of Synthetic Images

2023-05-25 · Xiaodan Xing, Federico Felder, Yang Nan, Giorgos Papanastasiou, Walsh Simon, Guang Yang

Synthetic images generated from deep generative models have the potential to address data scarcity and data privacy issues. The selection of synthesis models is mostly based on image quality measurements, and most researchers favor synthetic images that produce realistic images, i.e., images with good fidelity scores, such as low Fr\'echet Inception Distance (FID) and high Peak Signal-To-Noise Ratio (PSNR). However, the quality of synthetic images is not limited to fidelity, and a wide spectrum of metrics should be evaluated to comprehensively measure the quality of synthetic images. In addition, quality metrics are not truthful predictors of the utility of synthetic images, and the relations between these evaluation metrics are not yet clear. In this work, we have established a comprehensive set of evaluators for synthetic images, including fidelity, variety, privacy, and utility. By analyzing more than 100k chest X-ray images and their synthetic copies, we have demonstrated that there is an inevitable trade-off between synthetic image fidelity, variety, and privacy. In addition, we have empirically demonstrated that the utility score does not require images with both high fidelity and high variety. For intra- and cross-task data augmentation, mode-collapsed images and low-fidelity images can still demonstrate high utility. Finally, our experiments have also showed that it is possible to produce images with both high utility and privacy, which can provide a strong rationale for the use of deep generative models in privacy-preserving applications. Our study can shore up comprehensive guidance for the evaluation of synthetic images and elicit further developments for utility-aware deep generative models in medical image synthesis.

📄 PDF Abstract BibTeX arXiv:2305.18337

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationImage GenerationPrivacy Preserving

Similar Papers 제목 키워드 기반

Fundamental Limits of Perfect Concept Erasure

2025-03-25 · Somnath Basu Roy Chowdhury, Avinava Dubey, Ahmad Beirami, Rahul Kidambi 외

Concept erasure is the task of erasing information about a concept (e.g., gender or race) from a representation set while retaining the maximum possible utility -- information from original representations. Concept erasu…

Fairness

Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models

2024-06-18 · Jie Chen, Yupeng Zhang, Bingning Wang, Wayne Xin Zhao 외

Synthetic data has been proposed as a solution to address the issue of high-quality data scarcity in the training of large language models (LLMs). Studies have shown that synthetic data can effectively improve the perfor…

Instruction Following

Synthetic Data -- Anonymisation Groundhog Day

2020-11-13 · Theresa Stadler, Bristena Oprisanu, Carmela Troncoso

Synthetic data has been advertised as a silver-bullet solution to privacy-preserving data publishing that addresses the shortcomings of traditional anonymisation techniques. The promise is that synthetic data drawn from …

Privacy Preserving

How Transformers Solve Propositional Logic Problems: A Mechanistic Analysis

2024-11-06 · Guan Zhe Hong, Nishanth Dikkala, Enming Luo, Cyrus Rashtchian 외

Large language models (LLMs) have shown amazing performance on tasks that require planning and reasoning. Motivated by this, we investigate the internal mechanisms that underpin a network's ability to perform complex log…

Logical Reasoning

On the Generalization of Training-based ChatGPT Detection Methods

2023-10-02 · Han Xu, Jie Ren, Pengfei He, Shenglai Zeng 외

ChatGPT is one of the most popular language models which achieve amazing performance on various natural language tasks. Consequently, there is also an urgent need to detect the texts generated ChatGPT from human written.…