paper-with-me

Papers

General Bayesian Policy Learning

2026-02-27 · Masahiro Kato arxiv

This study proposes a General Bayes framework for policy learning. We consider decision problems in which a decision-maker chooses an action from a given set to maximize expected welfare. Typical examples include treatment choice and portfolio optimization. In such problems, the statistical target is a decision rule, and predicting each potential outcome is not necessarily of primary interest. We formulate this policy-learning problem through loss-based Bayesian updating. Our main technical device is a squared-loss surrogate for welfare maximization. We show that maximizing empirical welfare over a policy class with a quadratic penalty controlled by a tuning parameter $ζ>0$ is equivalent to minimizing a scaled squared error in the outcome difference. The resulting General Bayes posterior over decision rules admits two equivalent characterizations: a Gaussian pseudo-likelihood representation and a decision-theoretic loss-based characterization. As one implementation, we introduce GBPLNet, a neural network implementation with a tanh-squashed output. Finally, we establish PAC-Bayes-type guarantees for the surrogate risk and the corresponding penalized welfare.

📄 PDF Abstract BibTeX arXiv:2602.23672

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Factored Contextual Policy Search with Bayesian Optimization

2016-12-06 · Peter Karkus, Andras Kupcsik, David Hsu, Wee Sun Lee

Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different "contexts". Bayesian optimization approaches to contextual policy search…

Active LearningBayesian OptimizationPositionVocal Bursts Type Prediction

Deep Bayesian Quadrature Policy Optimization

2020-06-28 · Akella Ravi Tej, Kamyar Azizzadenesheli, Mohammad Ghavamzadeh, Anima Anandkumar 외

We study the problem of obtaining accurate policy gradient estimates using a finite number of samples. Monte-Carlo methods have been the default choice for policy gradient estimation, despite suffering from high variance…

continuous-controlContinuous ControlPolicy Gradient Methods

PAC-Bayesian Policy Evaluation for Reinforcement Learning

2012-02-14 · Mahdi Milani Fard, Joelle Pineau, Csaba Szepesvari

Bayesian priors offer a compact yet general means of incorporating domain knowledge into many learning tasks. The correctness of the Bayesian analysis and inference, however, largely depends on accuracy and correctness o…

Model Selectionreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling

2024-06-05 · Imad Aouali, Victor-Emmanuel Brunel, David Rohde, Anna Korba

Off-policy learning (OPL) often involves minimizing a risk estimator based on importance weighting to correct bias from the logging policy used to collect data. However, this method can produce an estimator with a high v…

Generalization Bounds

Factored Contextual Policy Search with Bayesian Optimization

2019-04-26 · Robert Pinsler, Peter Karkus, Andras Kupcsik, David Hsu 외

Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different task contexts. Contextual policy search offers data-efficient learning a…

Active LearningBayesian OptimizationPosition