paper-with-me

홈 › Papers

Convergence of Q-value in case of Gaussian rewards

2020-03-07 · Konatsu Miyamoto, Masaya Suzuki, Yuma Kigami, Kodai Satake

In this paper, as a study of reinforcement learning, we converge the Q function to unbounded rewards such as Gaussian distribution. From the central limit theorem, in some real-world applications it is natural to assume that rewards follow a Gaussian distribution , but existing proofs cannot guarantee convergence of the Q-function. Furthermore, in the distribution-type reinforcement learning and Bayesian reinforcement learning that have become popular in recent years, it is better to allow the reward to have a Gaussian distribution. Therefore, in this paper, we prove the convergence of the Q-function under the condition of $E[r(s,a)^2]<\infty$, which is much more relaxed than the existing research. Finally, as a bonus, a proof of the policy gradient theorem for distributed reinforcement learning is also posted.

📄 PDF Abstract BibTeX arXiv:2003.03526

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Stochastic differential equations for limiting description of UCB rule for Gaussian multi-armed bandits

2021-12-13 · Sergey Garbar

We consider the upper confidence bound strategy for Gaussian multi-armed bandits with known control horizon sizes $N$ and build its limiting description with a system of stochastic differential equations and ordinary dif…

Multi-Armed Bandits

Risk-Aware Algorithms for Combinatorial Semi-Bandits

2021-12-02 · Shaarad Ayyagari, Ambedkar Dukkipati

In this paper, we study the stochastic combinatorial multi-armed bandit problem under semi-bandit feedback. While much work has been done on algorithms that optimize the expected reward for linear as well as some general…

Reward-estimation variance elimination in sequential decision processes

2018-11-15 · Sergey Pankov

Policy gradient methods are very attractive in reinforcement learning due to their model-free nature and convergence guarantees. These methods, however, suffer from high variance in gradient estimation, resulting in poor…

Policy Gradient MethodsReinforcement Learning

Convergence Rates of Gaussian ODE Filters

2018-07-25 · Hans Kersting, T. J. Sullivan, Philipp Hennig

A recently-introduced class of probabilistic (uncertainty-aware) solvers for ordinary differential equations (ODEs) applies Gaussian (Kalman) filtering to initial value problems. These methods model the true solution $x$…

Convergence analysis of belief propagation for pairwise linear Gaussian models

2017-06-12 · Jian Du, Shaodan Ma, Yik-Chung Wu, Soummya Kar 외

Gaussian belief propagation (BP) has been widely used for distributed inference in large-scale networks such as the smart grid, sensor networks, and social networks, where local measurements/observations are scattered ov…