paper-with-me

홈 › Papers

Policy Gradient with Kernel Quadrature

2023-10-23 · Satoshi Hayakawa, Tetsuro Morimura

Reward evaluation of episodes becomes a bottleneck in a broad range of reinforcement learning tasks. Our aim in this paper is to select a small but representative subset of a large batch of episodes, only on which we actually compute rewards for more efficient policy gradient iterations. We build a Gaussian process modeling of discounted returns or rewards to derive a positive definite kernel on the space of episodes, run an ``episodic" kernel quadrature method to compress the information of sample episodes, and pass the reduced episodes to the policy network for gradient updates. We present the theoretical background of this procedure as well as its numerical illustrations in MuJoCo tasks.

📄 PDF Abstract BibTeX arXiv:2310.14768

Code (0)

등록된 구현이 없습니다.

Tasks

Causal DiscoveryMuJoCo

Methods 이 논문이 사용한 방법론

Gaussian Process Gaussian Processes are non-parametric models for approximating functions. They rely upon a measure of similarity between points (the kernel function) to predict the value for…

Similar Papers 제목 키워드 기반

Deep Bayesian Quadrature Policy Optimization

2020-06-28 · Akella Ravi Tej, Kamyar Azizzadenesheli, Mohammad Ghavamzadeh, Anima Anandkumar 외

We study the problem of obtaining accurate policy gradient estimates using a finite number of samples. Monte-Carlo methods have been the default choice for policy gradient estimation, despite suffering from high variance…

continuous-controlContinuous ControlPolicy Gradient Methods

Sparse solutions of the kernel herding algorithm by improved gradient approximation

2021-05-17 · Kazuma Tsuji, Ken'ichiro Tanaka

The kernel herding algorithm is used to construct quadrature rules in a reproducing kernel Hilbert space (RKHS). While the computational efficiency of the algorithm and stability of the output quadrature formulas are adv…

Computational Efficiency

Impact of Computation in Integral Reinforcement Learning for Continuous-Time Control

2024-02-27 · Wenhan Cao, Wei Pan

Integral reinforcement learning (IntRL) demands the precise computation of the utility function's integral at its policy evaluation (PEV) stage. This is achieved through quadrature rules, which are weighted sums of utili…

Kernel quadrature with DPPs

2019-06-18 · NeurIPS 2019 12 · Ayoub Belhadji, Rémi Bardenet, Pierre Chainais

We study quadrature rules for functions from an RKHS, using nodes sampled from a determinantal point process (DPP). DPPs are parametrized by a kernel, and we use a truncated and saturated version of the RKHS kernel. This…

Fast Approximation and Estimation Bounds of Kernel Quadrature for Infinitely Wide Models

2019-02-02 · Sho Sonoda

An infinitely wide model is a weighted integration $\int \varphi(x,v) d \mu(v)$ of feature maps. This model excels at handling an infinite number of features, and thus it has been adopted to the theoretical study of deep…

Model SelectionNumerical Integration