paper-with-me

Papers

Dead Directions: Geometric Singular Learning

2026-06-04 · Tejas Pradeep Shirodkar arxiv

Singular learning theory and information geometry study the same spaces: the former in resolved coordinates, the latter in original coordinates under a non-degeneracy assumption that overparameterised models violate. This paper carries one direction of the bridge between them, from Watanabe's invariants to Fisher geometry, through one primitive, the dead direction: a unit vector along which the Fisher metric degenerates, equivalently a direction crossing the analytic singular set along which the KL divergence keeps a zero of high order, its KL order set by how fast that divergence vanishes. Our central result recovers the KL order as the decay rate of the directional Fisher quadratic form approaching the singularity, in original coordinates, without a Hironaka resolution. A selection rule on smooth fibres translates this rate into Watanabe's single-direction contribution to the real log canonical threshold, and the recovery extends to multi-component crossings, multiplicity $m$, the singular fluctuation $ν$, prior-RLCT shifts, and tempered posteriors. We then carry the rate into a deep network: a multi-layer K-FAC factorisation writes each Fisher block as a product of activation- and gradient-side rates with a duality between them, instantiated at residual streams, layer normalisation, and attention. A quotient theorem carries the rate to the gauge quotient for optimizers whose update commutes with the group action; Adam's per-coordinate preconditioner fails that condition, so we construct DDCAdam, an equivariant Adam-family preconditioner, and prove the quotient rate along its trajectory. The result is a trajectory-rate readout of Watanabe's triple $(λ, m, ν)$ from one checkpoint's forward and backward passes, without posterior sampling.

📄 PDF Abstract BibTeX arXiv:2606.05957

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Algebraic Dead Directions in LayerNorm Transformers: A Forward-Pass-Only Diagnostic at LLM Scale

2026-06-17 · Tejas Pradeep Shirodkar, P. J. Narayanan arxiv

Pretrained transformers sit near singular minima of the loss, where the Fisher information metric degenerates along dead directions: directions in parameter space along which the directional Fisher vanishes. Locating suc…

Dead-Direction Signatures: A Cheap Spectral Reading of Singular Complexity

2026-06-19 · Tejas Pradeep Shirodkar, P. J. Narayanan arxiv

Singular learning theory characterises the complexity of a deep network through the geometry of its loss singularities. The local learning coefficient (LLC), the standard estimator of Watanabe's real log canonical thresh…

Measuring Dead Directions: Decomposing and Classifying Singular Structure off Canonical Alignment

2026-07-01 · Tejas Pradeep Shirodkar arxiv

We give a descent-free, alignment-free measurement of singular structure on trained networks. At a single frozen checkpoint the read recovers the order $k$ of each dead direction from the directional-Fisher rate, the mas…

TokenBlowUp: Resolving Representational Singularities in LLM Token Spaces via Monoidal Transformations

2025-07-26 · Dongfang Zhao arxiv

Recent work has provided compelling evidence challenging the foundational manifold hypothesis for the token embedding spaces of Large Language Models (LLMs). These findings reveal the presence of geometric singularities …

Singular Fluctuation as Specific Heat in Bayesian Learning

2025-12-24 · Sean Plummer arxiv

Singular learning theory characterizes Bayesian models with non-identifiable parameterizations through two central quantities: the real log canonical threshold (RLCT), which governs marginal likelihood asymptotics, and t…