paper-with-me

Papers

DOP: Deep Optimistic Planning with Approximate Value Function Evaluation

2018-03-22 · Francesco Riccio, Roberto Capobianco, Daniele Nardi

Research on reinforcement learning has demonstrated promising results in manifold applications and domains. Still, efficiently learning effective robot behaviors is very difficult, due to unstructured scenarios, high uncertainties, and large state dimensionality (e.g. multi-agent systems or hyper-redundant robots). To alleviate this problem, we present DOP, a deep model-based reinforcement learning algorithm, which exploits action values to both (1) guide the exploration of the state space and (2) plan effective policies. Specifically, we exploit deep neural networks to learn Q-functions that are used to attack the curse of dimensionality during a Monte-Carlo tree search. Our algorithm, in fact, constructs upper confidence bounds on the learned value function to select actions optimistically. We implement and evaluate DOP on different scenarios: (1) a cooperative navigation problem, (2) a fetching task for a 7-DOF KUKA robot, and (3) a human-robot handover with a humanoid robot (both in simulation and real). The obtained results show the effectiveness of DOP in the chosen applications, where action values drive the exploration and reduce the computational demand of the planning process while achieving good performance.

📄 PDF Abstract BibTeX arXiv:1803.08501

Code (0)

등록된 구현이 없습니다.

Tasks

Model-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Optimistic Planning by Regularized Dynamic Programming

2023-02-27 · Antoine Moulin, Gergely Neu

We propose a new method for optimistic planning in infinite-horizon discounted Markov decision processes based on the idea of adding regularization to the updates of an otherwise standard approximate value iteration proc…

Influence-Optimistic Local Values for Multiagent Planning --- Extended Version

2015-02-18 · Frans A. Oliehoek, Matthijs T. J. Spaan, Stefan Witwicki

Recent years have seen the development of methods for multiagent planning under uncertainty that scale to tens or even hundreds of agents. However, most of these methods either make restrictive assumptions on the problem…

BenchmarkingHeuristic Search

Learning to Plan via Deep Optimistic Value Exploration

2020-06-08 · L4DC 2020 6 · Tim Seyde, Wilko Schwarting, Sertac Karaman, Daniela Rus

Deep exploration requires coordinated long-term planning. We present a model-based reinforcement learning algorithm that guides policy learning through a value function that exhibits optimism in the face of uncertainty. …

BenchmarkingModel-based Reinforcement LearningReinforcement Learning (RL)

Randomized Exploration for Reinforcement Learning with General Value Function Approximation

2021-06-15 · Haque Ishfaq, Qiwen Cui, Viet Nguyen, Alex Ayoub 외

We propose a model-free reinforcement learning algorithm inspired by the popular randomized least squares value iteration (RLSVI) algorithm as well as the optimism principle. Unlike existing upper-confidence-bound (UCB) …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Optimistic policy iteration and natural actor-critic: A unifying view and a non-optimality result

2013-12-01 · NeurIPS 2013 12 · Paul Wagner

Approximate dynamic programming approaches to the reinforcement learning problem are often categorized into greedy value function methods and value-based policy gradient methods. As our first main result, we show that an…

Policy Gradient MethodsReinforcement Learning