paper-with-me

홈 › Papers

Marginals Before Conditionals

2026-03-10 · Mihir Sahasrabudhe arxiv

We construct a minimal task that isolates conditional learning in neural networks: a surjective map with K-fold ambiguity, resolved by a selector token z, so H(A | B) = log K while H(A | B, z) = 0. The model learns the marginal P(A | B) first, producing a plateau at exactly log K, before acquiring the full conditional in a sharp, collective transition. The plateau has a clean decomposition: height = log K (set by ambiguity), duration = f(D) (set by dataset size D, not K). Gradient noise stabilizes the marginal solution: higher learning rates monotonically slow the transition (3.6* across a 7* η range at fixed throughput), and batch-size reduction delays escape, consistent with an entropic force opposing departure from the low-gradient marginal. Internally, a selector-routing head assembles during the plateau, leading the loss transition by ~50% of the waiting time. This is the Type 2 directional asymmetry of Papadopoulos et al. [2024], measured dynamically: we track the excess risk from log K to zero and characterize what stabilizes it, what triggers its collapse, and how long it takes.

📄 PDF Abstract BibTeX arXiv:2603.10074

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

KERMIT: Generative Insertion-Based Modeling for Sequences

2019-06-04 · William Chan, Nikita Kitaev, Kelvin Guu, Mitchell Stern 외

We present KERMIT, a simple insertion-based approach to generative modeling for sequences and sequence pairs. KERMIT models the joint distribution and its decompositions (i.e., marginals and conditionals) using a single …

Machine TranslationQuestion AnsweringRepresentation LearningTranslation

Marginalizable Density Models

2021-06-08 · Dar Gilboa, Ari Pakman, Thibault Vatter

Probability density models based on deep networks have achieved remarkable success in modeling complex high-dimensional datasets. However, unlike kernel density estimators, modern neural models do not yield marginals or …

Density EstimationImputation

On The Chain Rule Optimal Transport Distance

2018-12-19 · Frank Nielsen, Ke Sun

We define a novel class of distances between statistical multivariate distributions by modeling an optimal transport problem on their marginals with respect to a ground distance defined on their conditionals. These new d…

Online Label Shift: Optimal Dynamic Regret meets Practical Algorithms

2023-09-21 · NeurIPS 2023 11

This paper focuses on supervised and unsupervised online label shift, where the class marginals $Q(y)$ varies but the class-conditionals $Q(x|y)$ remain invariant. In the unsupervised setting, our goal is to adapt a lear…

Bayesian Properties of Normalized Maximum Likelihood and its Fast Computation

2014-01-28 · Andrew Barron, Teemu Roos, Kazuho Watanabe

The normalized maximized likelihood (NML) provides the minimax regret solution in universal data compression, gambling, and prediction, and it plays an essential role in the minimum description length (MDL) method of sta…

Data CompressionPrediction