paper-with-me

홈 › Papers

SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data

2024-04-24 · Anton Danholt Lautrup, Tobias Hyrup, Arthur Zimek, Peter Schneider-Kamp

With the growing demand for synthetic data to address contemporary issues in machine learning, such as data scarcity, data fairness, and data privacy, having robust tools for assessing the utility and potential privacy risks of such data becomes crucial. SynthEval, a novel open-source evaluation framework distinguishes itself from existing tools by treating categorical and numerical attributes with equal care, without assuming any special kind of preprocessing steps. This~makes it applicable to virtually any synthetic dataset of tabular records. Our tool leverages statistical and machine learning techniques to comprehensively evaluate synthetic data fidelity and privacy-preserving integrity. SynthEval integrates a wide selection of metrics that can be used independently or in highly customisable benchmark configurations, and can easily be extended with additional metrics. In this paper, we describe SynthEval and illustrate its versatility with examples. The framework facilitates better benchmarking and more consistent comparisons of model capabilities.

📄 PDF Abstract BibTeX arXiv:2404.15821

Code (1)

schneiderkamplab/syntheval 공식 구현

Tasks

BenchmarkingFairnessPrivacy Preserving

Similar Papers 제목 키워드 기반

SYNTHEVAL: Hybrid Behavioral Testing of NLP Models with Synthetic CheckLists

2024-08-30 · Raoyuan Zhao, Abdullatif Köksal, Yihong Liu, Leonie Weissweiler 외

Traditional benchmarking in NLP typically involves using static held-out test sets. However, this approach often results in an overestimation of performance and lacks the ability to offer comprehensive, interpretable, an…

BenchmarkingSentiment Analysis

ConfusionPrompt: Practical Private Inference for Online Large Language Models

2023-12-30 · Peihua Mai, Youjia Yang, Ran Yan, Rui Ye 외

State-of-the-art large language models (LLMs) are typically deployed as online services, requiring users to transmit detailed prompts to cloud servers. This raises significant privacy concerns. In response, we introduce …

Privacy PreservingZero-shot Generalization

The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization

2026-01-17 · Natalia Tomashenko, Xiaoxiao Miao, Pierre Champion, Sarina Meyer 외 arxiv

We present results and analyses from the third VoicePrivacy Challenge held in 2024, which focuses on advancing voice anonymization technologies. The task was to develop a voice anonymization system for speech data that c…

Probably Approximately Correct Federated Learning

2023-04-10 · Xiaojin Zhang, Anbu Huang, Lixin Fan, Kai Chen 외

Federated learning (FL) is a new distributed learning paradigm, with privacy, utility, and efficiency as its primary pillars. Existing research indicates that it is unlikely to simultaneously attain infinitesimal privacy…

Federated LearningPAC learning

Robust Utility-Preserving Text Anonymization Based on Large Language Models

2024-07-16 · Tianyu Yang, Xiaodan Zhu, Iryna Gurevych

Text anonymization is crucial for sharing sensitive data while maintaining privacy. Existing techniques face the emerging challenges of re-identification attack ability of Large Language Models (LLMs), which have shown a…

Text Anonymization