paper-with-me

홈 › Papers

Value Function Approximations via Kernel Embeddings for No-Regret Reinforcement Learning

2020-11-16 · Sayak Ray Chowdhury, Rafael Oliveira

We consider the regret minimization problem in reinforcement learning (RL) in the episodic setting. In many real-world RL environments, the state and action spaces are continuous or very large. Existing approaches establish regret guarantees by either a low-dimensional representation of the stochastic transition model or an approximation of the $Q$-functions. However, the understanding of function approximation schemes for state-value functions largely remains missing. In this paper, we propose an online model-based RL algorithm, namely the CME-RL, that learns representations of transition distributions as embeddings in a reproducing kernel Hilbert space while carefully balancing the exploitation-exploration tradeoff. We demonstrate the efficiency of our algorithm by proving a frequentist (worst-case) regret bound that is of order $\tilde{O}\big(H\gamma_N\sqrt{N}\big)$\footnote{ $\tilde{O}(\cdot)$ hides only absolute constant and poly-logarithmic factors.}, where $H$ is the episode length, $N$ is the total number of time steps and $\gamma_N$ is an information theoretic quantity relating the effective dimension of the state-action feature space. Our method bypasses the need for estimating transition probabilities and applies to any domain on which kernels can be defined. It also brings new insights into the general theory of kernel methods for approximate inference and RL regret minimization.

📄 PDF Abstract BibTeX arXiv:2011.07881

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Consequences of Kernel Regularity for Bandit Optimization

2025-12-05 · Madison Lee, Tara Javidi arxiv

In this work we investigate the relationship between kernel regularity and algorithmic performance in the bandit optimization of RKHS functions. While reproducing kernel Hilbert space (RKHS) methods traditionally rely on…

Provably Efficient Reinforcement Learning with Kernel and Neural Function Approximations

2020-12-01 · NeurIPS 2020 12 · Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang 외

Reinforcement learning (RL) algorithms combined with modern function approximators such as kernel functions and deep neural networks have achieved significant empirical successes in large-scale application problems with…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Dynamic Regret for Online Regression in RKHS via Discounted VAW and Subspace Approximation

2026-04-27 · Dmitry B. Rokhlin, Georgiy A. Karapetyants arxiv

We study online regression with the square loss in a reproducing kernel Hilbert space under a dynamic regret criterion. The learner is compared with a time-varying comparator sequence, and the bounds depend on its path l…

Kernelized Reinforcement Learning with Order Optimal Regret Bounds

2023-06-13 · NeurIPS 2023 11 · Sattar Vakili, Julia Olkhovskaya

Reinforcement learning (RL) has shown empirical success in various real world settings with complex models and large state-action spaces. The existing analytical results, however, typically focus on settings with a small…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Optimistic Policy Optimization with General Function Approximations

2021-01-01 · Qi Cai, Zhuoran Yang, Csaba Szepesvari, Zhaoran Wang

Although policy optimization with neural networks has a track record of achieving state-of-the-art results in reinforcement learning on various domains, the theoretical understanding of the computational and sample effic…

reinforcement-learningReinforcement Learning (RL)