paper-with-me

홈 › Papers

Shaping the Prior: How Synthetic Task Distributions Determine Tabular Foundation Model Quality

2026-05-18 · Mohamed Bouadi, Nassim Bouarour, Varun Kulkarni, Shivam Dubey, Aditya Tanna, Vinay Kumar Sankarapu arxiv

What determines the quality of a tabular foundation model? Unlike language or vision, tabular foundation models acquire their inductive biases almost entirely from synthetic pretraining distributions, yet the design of these distributions remains poorly understood. Standard synthetic priors are too well-behaved: they omit the irregularities and failure modes that determine deployment robustness. We introduce O'Prior, a compositional realism prior built around four coupled components: a hierarchical SCM meta-generator spanning diverse functional families; a modular realism engine covering heterogeneous marginals, missingness, and target transforms; an explicit stress module injecting confounding and support-query mismatch; and a curriculum-governed, leakage-safe generation protocol. To isolate prior design as the scientific variable, we hold architecture, optimizer, and compute budget fixed and vary only the synthetic task distribution. O'Prior yields consistent and substantial improvements in downstream accuracy and robustness across real tabular benchmarks, with gains concentrated in regimes characterized by distributional irregularities. Ablations confirm that mechanism diversity, realism composition, and shift-aware stress each contribute independently, their effects are not interchangeable. These results establish synthetic prior construction as a first-order and largely overlooked determinant of tabular foundation model quality

📄 PDF Abstract BibTeX arXiv:2605.18971

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Joint Learning of Probabilistic and Geometric Shaping for Coded Modulation Systems

2020-04-10 · Fayçal Ait Aoudia, Jakob Hoydis

We introduce a trainable coded modulation scheme that enables joint optimization of the bit-wise mutual information (BMI) through probabilistic shaping, geometric shaping, bit labeling, and demapping for a specific chann…

Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs

2025-10-01 · Leyla Mirvakhabova, Babak Ehteshami Bejnordi, Gaurav Kumar, Hanxue Liang 외 arxiv

Upcycling pre-trained dense models into sparse Mixture-of-Experts (MoEs) efficiently increases model capacity but often suffers from poor expert specialization due to naive weight replication. Our analysis reveals that u…

Partial Enumerative Sphere Shaping

2019-11-29

The dependency between the Gaussianity of the input distribution for the additive white Gaussian noise (AWGN) channel and the gap-to-capacity is discussed. We show that a set of particular approximations to the Maxwell-B…

Steering Language Generation: Harnessing Contrastive Expert Guidance and Negative Prompting for Coherent and Diverse Synthetic Data Generation

2023-08-15 · Charles O'Neill, Yuan-Sen Ting, Ioana Ciuca, Jack Miller 외

Large Language Models (LLMs) hold immense potential to generate synthetic data of high quality and utility, which has numerous applications from downstream model training to practical data utilisation. However, contempor…

Comment GenerationDiversitySynthetic Data GenerationText Generation

Reward Shaping via Meta-Learning

2019-01-27 · Haosheng Zou, Tongzheng Ren, Dong Yan, Hang Su 외

Reward shaping is one of the most effective methods to tackle the crucial yet challenging problem of credit assignment in Reinforcement Learning (RL). However, designing shaping functions usually requires much expert kno…

Meta-LearningReinforcement LearningReinforcement Learning (RL)