paper-with-me

Papers

Understanding Latent Flow Models for Tabular Data Synthesis: Targets, Paths, and Sampling

2026-06-18 · Bahrul Ilmi Nasution arxiv

Synthetic tabular data enables microdata sharing in regulated domains, yet deploying continuous-time generative models requires balancing analytical utility, disclosure risk, and computational cost. Latent-space flow models are flexible, but theoretical equivalences across learning targets, probability paths, and sampling dynamics can translate into different behaviour under finite-step integration and explicit compute budgets. We present an empirical study of tabular latent flow models across seven datasets, evaluating velocity, score, noise, and posterior matching objectives under optimal transport (OT) and variance-preserving (VP) paths, ODE and SDE sampling, and varying integration budgets. Our contributions are threefold: (1) we show that the learning target largely determines the utility-risk operating regime, with velocity and posterior matching tending to yield higher utility, while score and noise matching tend to achieve lower disclosure risk; (2) we demonstrate that configuration and sampling choices shift performance, with midpoint often improving distributional fidelity and OT paths often tolerating earlier stopping than VP, enabling compute savings under fixed budgets or risk thresholds; and (3) we distil these findings into actionable defaults and practical configuration guidance to support pre-release model selection under disclosure risk and resource constraints. The code implementation and supplementary materials can be accessed in https://github.com/rulnasution/tabular-latent-flow/.

📄 PDF Abstract BibTeX arXiv:2606.20878

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BSTabDiff: Block-Subunit Diffusion Priors for High-Dimensional Tabular Data Generation

2026-06-08 · Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali, Gianfranco Doretto 외 arxiv

High-Dimensional Low-Sample Size (HDLSS) tabular domains (e.g., omics) are characterized by $n \ll m$, where $n$ = number of samples, and $m$ = number of features. Such domains often exhibit strong local correlation grou…

Tabular Data Generation

Multimodal synthesis of MRI and tabular data with diffusion in a joint latent space via cross-attention

2026-05-05 · Daniel Mensing, Jan Kapar, Jochen G. Hirsch, Matthias Günther 외 arxiv

We propose a multimodal latent diffusion model that jointly synthesizes volumetric magnetic resonance imaging (MRI) and tabular clinical data within a shared latent space via cross-attention. This approach enables cohere…

Representation LearningImage Generation

Flow Matching for Tabular Data Synthesis

2025-11-30 · Bahrul Ilmi Nasution, Floor Eijkelboom, Mark Elliot, Richard Allmendinger 외 arxiv

Synthetic data generation is an important tool for privacy-preserving data sharing. Although diffusion models have set recent benchmarks, flow matching (FM) offers a promising alternative. This paper presents different w…

Synthetic Data Generation

A Systematic Framework for Tabular Data Disentanglement

2026-04-09 · Ivan Tjuawinata, Andre Gunawan, Anh Quan Tran, Nitish Kumar 외 arxiv

Tabular data, widely used in various applications such as industrial control systems, finance, and supply chain, often contains complex interrelationships among its attributes. Data disentanglement seeks to transform suc…

Tabular Data Generation

Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent Space

2023-10-14 · Hengrui Zhang, Jiani Zhang, Balasubramaniam Srinivasan, Zhengyuan Shen 외

Recent advances in tabular data generation have greatly enhanced synthetic data quality. However, extending diffusion models to tabular data is challenging due to the intricately varied distributions and a blend of data …

Tabular Data Generation