paper-with-me

홈 › Papers

In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning

2025-10-13 · Tomoya Wakayama, Taiji Suzuki arxiv

This paper develops a finite-sample statistical theory for in-context learning (ICL), analyzed within a meta-learning framework that accommodates mixtures of diverse task types. We introduce a principled risk decomposition that separates the total ICL risk into two orthogonal components: Bayes Gap and Posterior Variance. The Bayes Gap quantifies how well the trained model approximates the Bayes-optimal in-context predictor. For a uniform-attention Transformer, we derive a non-asymptotic upper bound on this gap, which explicitly clarifies the dependence on the number of pretraining prompts and their context length. The Posterior Variance is a model-independent risk representing the intrinsic task uncertainty. Our key finding is that this term is determined solely by the difficulty of the true underlying task, while the uncertainty arising from the task mixture vanishes exponentially fast with only a few in-context examples. Together, these results provide a unified view of ICL: the Transformer selects the optimal meta-algorithm during pretraining and rapidly converges to the optimal algorithm for the true task at test time.

📄 PDF Abstract BibTeX arXiv:2510.10981

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian Inference

Similar Papers 제목 키워드 기반

Learning-to-Optimize with PAC-Bayesian Guarantees: Theoretical Considerations and Practical Implementation

2024-04-04 · Michael Sucker, Jalal Fadili, Peter Ochs

We use the PAC-Bayesian theory for the setting of learning-to-optimize. To the best of our knowledge, we present the first framework to learn optimization algorithms with provable generalization guarantees (PAC-Bayesian …

Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling

2025-12-22 · Indranil Halder, Cengiz Pehlevan arxiv

Recent developments in large language models have shown advantages in reallocating a notable share of computational resource from training time to inference time. However, the principles behind inference time scaling are…

Analogy as Nonparametric Bayesian Inference over Relational Systems

2020-06-07 · Ruairidh M. Battleday, Thomas L. Griffiths

Much of human learning and inference can be framed within the computational problem of relational generalization. In this project, we propose a Bayesian model that generalizes relational knowledge to novel environments b…

Analogical SimilarityBayesian Inference

Excess risk analysis for epistemic uncertainty with application to variational inference

2022-06-02 · Futoshi Futami, Tomoharu Iwata, Naonori Ueda, Issei Sato 외

Bayesian deep learning plays an important role especially for its ability evaluating epistemic uncertainty (EU). Due to computational complexity issues, approximation methods such as variational inference (VI) have been …

Bayesian InferenceVariational Inference

Domain Agnostic Conditional Invariant Predictions for Domain Generalization

2024-06-09 · Zongbin Wang, Bin Pan, Zhenwei Shi

Domain generalization aims to develop a model that can perform well on unseen target domains by learning from multiple source domains. However, recent-proposed domain generalization models usually rely on domain labels, …

Bayesian InferenceDomain Generalization