paper-with-me

홈 › Papers

Critical Challenges and Guidelines in Evaluating Synthetic Tabular Data: A Systematic Review

2025-04-10 · Nazia Nafis, Inaki Esnaola, Alvaro Martinez-Perez, Maria-Cruz Villa-Uriol, Venet Osmani

Generating synthetic tabular data can be challenging, however evaluation of their quality is just as challenging, if not more. This systematic review sheds light on the critical importance of rigorous evaluation of synthetic health data to ensure reliability, relevance, and their appropriate use. Based on screening of 1766 papers and a detailed review of 101 papers we identified key challenges, including lack of consensus on evaluation methods, improper use of evaluation metrics, limited input from domain experts, inadequate reporting of dataset characteristics, and limited reproducibility of results. In response, we provide several guidelines on the generation and evaluation of synthetic data, to allow the community to unlock and fully harness the transformative potential of synthetic data and accelerate innovation.

📄 PDF Abstract BibTeX arXiv:2504.18544

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Improve Fidelity and Utility of Synthetic Credit Card Transaction Time Series from Data-centric Perspective

2024-01-01 · Din-Yin Hsieh, Chi-Hua Wang, Guang Cheng

Exploring generative model training for synthetic tabular data, specifically in sequential contexts such as credit card transaction data, presents significant challenges. This paper addresses these challenges, focusing o…

Fraud DetectionTime Series

SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data

2025-11-17 · Ke Yu, Shigeru Ishikura, Yukari Usukura, Yuki Shigoku 외 arxiv

Synthetic tabular data, which are widely used in domains such as healthcare, enterprise operations, and customer analytics, are increasingly evaluated to ensure that they preserve both privacy and utility. While existing…

Feature Importance

LLM-TabFlow: Synthetic Tabular Data Generation with Inter-column Logical Relationship Preservation

2025-03-04 · Yunbo Long, Liming Xu, Alexandra Brintrup

Synthetic tabular data have widespread applications in industrial domains such as healthcare, finance, and supply chains, owing to their potential to protect privacy and mitigate data scarcity. However, generating realis…

Large Language ModelTabular Data Generation

FedTabDiff: Federated Learning of Diffusion Probabilistic Models for Synthetic Mixed-Type Tabular Data Generation

2024-01-11 · Timur Sattarov, Marco Schreyer, Damian Borth

Realistic synthetic tabular data generation encounters significant challenges in preserving privacy, especially when dealing with sensitive information in domains like finance and healthcare. In this paper, we introduce …

AttributeDenoisingFederated LearningTabular Data Generation

ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control

2025-10-21 · Yuzheng Hu, Ryan McKenna, Da Yu, Shanshan Wu 외 arxiv

Generating high-quality synthetic text under differential privacy (DP) is critical for training and evaluating language models without compromising user privacy. Prior work on synthesizing DP datasets often fail to prese…

Conditional Text Generation