BSTabDiff: Block-Subunit Diffusion Priors for High-Dimensional Tabular Data Generation
High-Dimensional Low-Sample Size (HDLSS) tabular domains (e.g., omics) are characterized by $n \ll m$, where $n$ = number of samples, and $m$ = number of features. Such domains often exhibit strong local correlation groups, sparse cross-group dependencies, heavy-tailed non-Gaussian marginals, heteroscedastic noise, and structured missingness, making direct density learning in $\mathbb{R}^m$ ill-conditioned since $n \ll m$. We propose BSTabDiff, a block-subunit generative framework that partitions the $m$ observed features into $M$ latent blocks ($M \ll m$) and generates each block via a shared low-dimensional subunit variable, concentrating global dependence learning in the compact block-latent space $\mathbb{R}^M$ while decoding to the full feature space with copula-driven dependence, flexible per-feature marginals, and explicit missingness mechanisms. BSTabDiff supports modern deep priors on block latents, including diffusion and normalizing flows, enabling stable synthesis and controllable benchmark generation in the HDLSS regime. Empirically, BSTabDiff produces more realistic and stable high-dimensional synthetic data when compared with unstructured tabular generators on HDLSS data.
Code (0)
등록된 구현이 없습니다.
Tasks
Tabular Data GenerationSimilar Papers 제목 키워드 기반
Analyzing conformational changes in single FRET-labeled A1 parts of archaeal A1AO-ATP synthase
ATP synthases utilize a proton motive force to synthesize ATP. In reverse, these membrane-embedded enzymes can also hydrolyze ATP to pump protons over the membrane. To prevent wasteful ATP hydrolysis, distinct control me…
Simba: Mamba augmented U-ShiftGCN for Skeletal Action Recognition in Videos
Skeleton Action Recognition (SAR) involves identifying human actions using skeletal joint coordinates and their interconnections. While plain Transformers have been attempted for this task, they still fall short compared…
Action RecognitionAction Recognition In VideosMambaTraining-Free Large Model Priors for Multiple-in-One Image Restoration
Image restoration aims to reconstruct the latent clear images from their degraded versions. Despite the notable achievement, existing methods predominantly focus on handling specific degradation types and thus require sp…
Image RestorationDiffusion In Diffusion: Reclaiming Global Coherence in Semi-Autoregressive Diffusion
One of the most compelling features of global discrete diffusion language models is their global bidirectional contextual capability. However, existing block-based diffusion studies tend to introduce autoregressive prior…
Polylogarithmic equilibrium treatment of molecular aggregation and critical concentrations
A full equilibrium treatment of molecular aggregation is presented for prototypes of 1D and 3D aggregates, with and without nucleation. By skipping complex kinetic parameters like aggregate size-dependent diffusion, the …