paper-with-me

홈 › Papers

TabQueryBench: A Query-Centric Benchmark for Synthetic Tabular Data

2026-07-04 · Jialin Zhang, Fenghao Dong, Yajie Zhou, Vyas Sekar, Shinan Liu arxiv

Synthetic tabular data support use cases like data sharing, model development under access restrictions, and rapid prototyping of analytical workflows. Modern generative models are evaluated by their statistical similarity, correlation structure, privacy, and downstream machine-learning utility. However, such evaluations leave a gap: they rarely test the structure that matters for analytical queries. We present TabQueryBench, a query-centric benchmark that uses SQL-shaped analytical queries as structural assessors for synthetic data fidelity. It provides an extensible foundation for query-centric synthetic-data evaluation. From 12 public sources of analytical queries, TabQueryBench taxonomizes recurring cross-domain logic into 44 reusable query templates and grounds them to each dataset via a policy-guided template-to-SQL pipeline. This makes queries schema-aware while preserving comparability across generative models. Across 49 datasets and 11 generative models, it activates 10-12 templates per dataset, producing more than 100 executable SQL queries per dataset. Our systematic experiments show five main patterns. First, current tabular generative models can have good distance-based fidelity, but they still fall short on query-centric fidelity: RealTabFormer achieves the highest query-centric fidelity, but it only reaches 0.75 +/- 0.15 (REAL data score is 1.00). Second, tabular generative models struggle with very high-cardinality discrete support. Third, SOTA generative models preserve good global conditional query-centric fidelity, but fail more on local queries. Fourth, tail fidelity deteriorates as queries move toward the extreme tail; even the best model recovers only about 40.7% of real rare values. Finally, there is a fidelity-cost tradeoff in tabular generation: BayesNet offers the strongest tradeoff, with slightly lower query-centric fidelity but much lower generation cost.

📄 PDF Abstract BibTeX arXiv:2607.03926

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reimagining Synthetic Tabular Data Generation through Data-Centric AI: A Comprehensive Benchmark

2023-10-25 · NeurIPS 2023 11 · Lasse Hansen, Nabeel Seedat, Mihaela van der Schaar, Andrija Petrovic

Synthetic data serves as an alternative in training machine learning models, particularly when real-world data is limited or inaccessible. However, ensuring that synthetic data mirrors the complex nuances of real-world d…

feature selectionModel SelectionSynthetic Data GenerationTabular Data Generation

SOMTab: Set-Order Mamba for Efficient Tabular In-Context Learning

2026-08-28 · Hao Wang, Siyu Zhang, Wei Ma arxiv

Tabular foundation models based on in-context learning have recently emerged as strong alternatives to task-specific model fitting. However, the current performance frontier remains dominated by attention-heavy architect…

Shaping the Prior: How Synthetic Task Distributions Determine Tabular Foundation Model Quality

2026-05-18 · Mohamed Bouadi, Nassim Bouarour, Varun Kulkarni, Shivam Dubey 외 arxiv

What determines the quality of a tabular foundation model? Unlike language or vision, tabular foundation models acquire their inductive biases almost entirely from synthetic pretraining distributions, yet the design of t…

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

2026-01-30 · Georgi Ganev, Emiliano De Cristofaro arxiv

Training generative machine learning models to produce synthetic tabular data has become a popular approach for enhancing privacy in data sharing. As this typically involves processing sensitive personal information, rel…

Synthetic Data Generation

Towards High Supervised Learning Utility Training Data Generation: Data Pruning and Column Reordering

2025-07-14 · Tung Sum Thomas Kwok, Zeyong Zhang, Chi-Hua Wang, Guang Cheng arxiv

Tabular data synthesis for supervised learning ('SL') model training is gaining popularity in industries such as healthcare, finance, and retail. Despite the progress made in tabular data generators, models trained with …