paper-with-me

홈 › Papers

Geometric Scaling of Bayesian Inference in LLMs

2025-12-27 · Naman Agarwal, Siddhartha R. Dalal, Vishal Misra arxiv

Recent work has shown that small transformers trained in controlled "wind-tunnel'' settings can implement exact Bayesian inference, and that their training dynamics produce a geometric substrate -- low-dimensional value manifolds and progressively orthogonal keys -- that encodes posterior structure. We investigate whether this geometric signature persists in production-grade language models. Across Pythia, Phi-2, Llama-3, and Mistral families, we find that last-layer value representations organize along a single dominant axis whose position strongly correlates with predictive entropy, and that domain-restricted prompts collapse this structure into the same low-dimensional manifolds observed in synthetic settings. To probe the role of this geometry, we perform targeted interventions on the entropy-aligned axis of Pythia-410M during in-context learning. Removing or perturbing this axis selectively disrupts the local uncertainty geometry, whereas matched random-axis interventions leave it intact. However, these single-layer manipulations do not produce proportionally specific degradation in Bayesian-like behavior, indicating that the geometry is a privileged readout of uncertainty rather than a singular computational bottleneck. Taken together, our results show that modern language models preserve the geometric substrate that enables Bayesian inference in wind tunnels, and organize their approximate Bayesian updates along this substrate.

📄 PDF Abstract BibTeX arXiv:2512.23752

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian Inference

Similar Papers 제목 키워드 기반

Variational Combinatorial Sequential Monte Carlo for Bayesian Phylogenetics in Hyperbolic Space

2025-01-29 · Alex Chen, Philipe Chlenski, Kenneth Munyuza, Antonio Khalil Moretti 외

Hyperbolic space naturally encodes hierarchical structures such as phylogenies (binary trees), where inward-bending geodesics reflect paths through least common ancestors, and the exponential growth of neighborhoods mirr…

Variational Inference

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space

2026-05-12 · Eric Bigelow, Raphaël Sarfati, Daniel Wurgaft, Owen Lewis 외 arxiv

Large Language Models (LLMs) update their behavior in context, which can be viewed as a form of Bayesian inference. However, the structure of the latent hypothesis space over which this inference operates remains unclear…

Bayesian Inference

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

2026-06-29 · Ankur Samanta, Akshayaa Magesh, Tal Lancewicki, Ayush Jain 외 arxiv

Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environment. Acting rationally then requires inf…

Large Scale Nonparametric Bayesian Inference: Data Parallelisation in the Indian Buffet Process

2009-12-01 · NeurIPS 2009 12 · Finale Doshi-Velez, Shakir Mohamed, Zoubin Ghahramani, David A. Knowles

Nonparametric Bayesian models provide a framework for flexible probabilistic modelling of complex datasets. Unfortunately, Bayesian inference methods often require high-dimensional averages and can be slow to compute, es…

Bayesian Inference

Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective

2024-06-02 · Fabian Falck, Ziyu Wang, Chris Holmes

In-context learning (ICL) has emerged as a particularly remarkable characteristic of Large Language Models (LLM): given a pretrained LLM and an observed dataset, LLMs can make predictions for new data points from the sam…

Bayesian InferenceIn-Context Learning