paper-with-me

Papers

Finite-time Analysis for the Knowledge-Gradient Policy

2016-06-15 · Yingfei Wang, Warren Powell

We consider sequential decision problems in which we adaptively choose one of finitely many alternatives and observe a stochastic reward. We offer a new perspective of interpreting Bayesian ranking and selection problems as adaptive stochastic multi-set maximization problems and derive the first finite-time bound of the knowledge-gradient policy for adaptive submodular objective functions. In addition, we introduce the concept of prior-optimality and provide another insight into the performance of the knowledge gradient policy based on the submodular assumption on the value of information. We demonstrate submodularity for the two-alternative case and provide other conditions for more general problems, bringing out the issue and importance of submodularity in learning problems. Empirical experiments are conducted to further illustrate the finite time behavior of the knowledge gradient policy.

📄 PDF Abstract BibTeX arXiv:1606.04624

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Finite-Time Analysis of Two Time-Scale Actor-Critic Methods

2020-12-01 · NeurIPS 2020 12 · Yue Wu, Weitong Zhang, Pan Xu, Quanquan Gu

Actor-critic (AC) methods have exhibited great empirical success compared with other reinforcement learning algorithms, where the actor uses the policy gradient to improve the learning policy and the critic uses temporal…

Vocal Bursts Valence Prediction

A Finite Time Analysis of Two Time-Scale Actor Critic Methods

2020-05-04 · Yue Wu, Weitong Zhang, Pan Xu, Quanquan Gu

Actor-critic (AC) methods have exhibited great empirical success compared with other reinforcement learning algorithms, where the actor uses the policy gradient to improve the learning policy and the critic uses temporal…

Vocal Bursts Valence Prediction

Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods

2021-09-13 · Xin Guo, Anran Hu, Junzi Zhang

When designing algorithms for finite-time-horizon episodic reinforcement learning problems, a common approach is to introduce a fictitious discount factor and use stationary policies for approximations. Empirically, it h…

Policy Gradient Methodsreinforcement-learningReinforcement LearningReinforcement Learning (RL)

On Linear Convergence of Policy Gradient Methods for Finite MDPs

2020-07-21 · Jalaj Bhandari, Daniel Russo

We revisit the finite time analysis of policy gradient methods in the one of the simplest settings: finite state and action MDPs with a policy class consisting of all stochastic policies and with exact gradient evaluatio…

Policy Gradient Methods

Finite-Sample Analysis of Proximal Gradient TD Algorithms

2020-06-06 · Bo Liu, Ji Liu, Mohammad Ghavamzadeh, Sridhar Mahadevan 외

In this paper, we analyze the convergence rate of the gradient temporal difference learning (GTD) family of algorithms. Previous analyses of this class of algorithms use ODE techniques to prove asymptotic convergence, an…