paper-with-me

홈 › Papers

CuTS: Customizable Tabular Synthetic Data Generation

2023-07-07 · Mark Vero, Mislav Balunović, Martin Vechev

Privacy, data quality, and data sharing concerns pose a key limitation for tabular data applications. While generating synthetic data resembling the original distribution addresses some of these issues, most applications would benefit from additional customization on the generated data. However, existing synthetic data approaches are limited to particular constraints, e.g., differential privacy (DP) or fairness. In this work, we introduce CuTS, the first customizable synthetic tabular data generation framework. Customization in CuTS is achieved via declarative statistical and logical expressions, supporting a wide range of requirements (e.g., DP or fairness, among others). To ensure high synthetic data quality in the presence of custom specifications, CuTS is pre-trained on the original dataset and fine-tuned on a differentiable loss automatically derived from the provided specifications using novel relaxations. We evaluate CuTS over four datasets and on numerous custom specifications, outperforming state-of-the-art specialized approaches on several tasks while being more general. In particular, at the same fairness level, we achieve 2.3% higher downstream accuracy than the state-of-the-art in fair synthetic data generation on the Adult dataset.

📄 PDF Abstract BibTeX arXiv:2307.03577

Code (1)

eth-sri/cuts 공식 구현 pytorch

Tasks

FairnessSynthetic Data GenerationTabular Data Generation

Similar Papers 제목 키워드 기반

dpmm: Differentially Private Marginal Models, a Library for Synthetic Tabular Data Generation

2025-05-31 · Sofiane Mahiou, Amir Dizche, Reza Nazari, Xinmin Wu 외

We propose dpmm, an open-source library for synthetic data generation with Differentially Private (DP) guarantees. It includes three popular marginal models -- PrivBayes, MST, and AIM -- that achieve superior utility and…

Synthetic Data GenerationTabular Data Generation

Preserving logical and functional dependencies in synthetic tabular data

2024-09-26 · Chaithra Umesh, Kristian Schultz, Manjunath Mahendra, Saparshi Bej 외

Dependencies among attributes are a common aspect of tabular data. However, whether existing tabular data generation algorithms preserve these dependencies while generating synthetic data is yet to be explored. In additi…

AttributeSynthetic Data GenerationTabular Data Generation

Privacy-Preserving Tabular Synthetic Data Generation Using TabularARGN

2025-08-08 · Andrey Sidorenko, Paul Tiwald arxiv

Synthetic data generation has become essential for securely sharing and analyzing sensitive data sets. Traditional anonymization techniques, however, often fail to adequately preserve privacy. We introduce the Tabular Au…

Synthetic Data Generation

Hierarchical Conditional Tabular GAN for Multi-Tabular Synthetic Data Generation

2024-11-11 · Wilhelm Ågren, Victorio Úbeda Sosa

The generation of synthetic data is a state-of-the-art approach to leverage when access to real data is limited or privacy regulations limit the usability of sensitive data. A fair amount of research has been conducted o…

Synthetic Data Generation

Generating Synthetic Relational Tabular Data via Structural Causal Models

2025-07-04 · Frederik Hoppe, Astrid Franz, Lars Kleinemeier, Udo Göbel arxiv

Synthetic tabular data generation has received increasing attention in recent years, particularly with the emergence of foundation models for tabular data. The breakthrough success of TabPFN (Hollmann et al.,2025), which…

Tabular Data Generation