paper-with-me

Papers

Manifold Regularization for Kernelized LSTD

2017-10-15 · Xinyan Yan, Krzysztof Choromanski, Byron Boots, Vikas Sindhwani

Policy evaluation or value function or Q-function approximation is a key procedure in reinforcement learning (RL). It is a necessary component of policy iteration and can be used for variance reduction in policy gradient methods. Therefore its quality has a significant impact on most RL algorithms. Motivated by manifold regularized learning, we propose a novel kernelized policy evaluation method that takes advantage of the intrinsic geometry of the state space learned from data, in order to achieve better sample efficiency and higher accuracy in Q-function approximation. Applying the proposed method in the Least-Squares Policy Iteration (LSPI) framework, we observe superior performance compared to widely used parametric basis functions on two standard benchmarks in terms of policy quality.

📄 PDF Abstract BibTeX arXiv:1710.05387

Code (0)

등록된 구현이 없습니다.

Tasks

Policy Gradient MethodsReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Properties of the Least Squares Temporal Difference learning algorithm

2013-01-22 · Kamil Ciosek

This paper presents four different ways of looking at the well-known Least Squares Temporal Differences (LSTD) algorithm for computing the value function of a Markov Reward Process, each of them leading to different insi…

LSTD: A Low-Shot Transfer Detector for Object Detection

2018-03-05 · Hao Chen, Yali Wang, Guoyou Wang, Yu Qiao

Recent advances in object detection are mainly driven by deep learning with large-scale detection benchmarks. However, the fully-annotated training set is often limited for a target detection task, which may deteriorate …

Few-Shot Object DetectionObjectobject-detectionObject Detection+1

On the Convergence of Gradient Descent in GANs: MMD GAN As a Gradient Flow

2020-11-04 · Youssef Mroueh, Truyen Nguyen

We consider the maximum mean discrepancy ($\mathrm{MMD}$) GAN problem and propose a parametric kernelized gradient flow that mimics the min-max game in gradient regularized $\mathrm{MMD}$ GAN. We show that this flow prov…

Finite Sample Analysis of LSTD with Random Projections and Eligibility Traces

2018-05-25 · Haifang Li, Yingce Xia, Wensheng Zhang

Policy evaluation with linear function approximation is an important problem in reinforcement learning. When facing high-dimensional feature spaces, such a problem becomes extremely hard considering the computation effic…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Predicting Multi-Antenna Frequency-Selective Channels via Meta-Learned Linear Filters based on Long-Short Term Channel Decomposition

2022-03-23 · Sangwoo Park, Osvaldo Simeone

An efficient data-driven prediction strategy for multi-antenna frequency-selective channels must operate based on a small number of pilot symbols. This paper proposes novel channel prediction algorithms that address this…

Meta-LearningPrediction