paper-with-me

Papers

Difference of Convex Functions Programming Applied to Control with Expert Data

2016-06-03 · Bilal Piot, Matthieu Geist, Olivier Pietquin

This paper reports applications of Difference of Convex functions (DC) programming to Learning from Demonstrations (LfD) and Reinforcement Learning (RL) with expert data. This is made possible because the norm of the Optimal Bellman Residual (OBR), which is at the heart of many RL and LfD algorithms, is DC. Improvement in performance is demonstrated on two specific algorithms, namely Reward-regularized Classification for Apprenticeship Learning (RCAL) and Reinforcement Learning with Expert Demonstrations (RLED), through experiments on generic Markov Decision Processes (MDP), called Garnets.

📄 PDF Abstract BibTeX arXiv:1606.01128

Code (0)

등록된 구현이 없습니다.

Tasks

General Classificationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Difference of Convex Functions Programming for Reinforcement Learning

2014-12-01 · NeurIPS 2014 12 · Bilal Piot, Matthieu Geist, Olivier Pietquin

Large Markov Decision Processes (MDPs) are usually solved using Approximate Dynamic Programming (ADP) methods such as Approximate Value Iteration (AVI) or Approximate Policy Iteration (API). The main contribution of this…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Piecewise Linear Regression via a Difference of Convex Functions

2020-07-05 · ICML 2020 1 · Ali Siahkamari, Aditya Gangrade, Brian Kulis, Venkatesh Saligrama

We present a new piecewise linear regression methodology that utilizes fitting a difference of convex functions (DC functions) to the data. These are functions $f$ that may be represented as the difference $\phi_1 - \phi…

regression

Convex Programs and Lyapunov Functions for Reinforcement Learning: A Unified Perspective on the Analysis of Value-Based Methods

2022-02-14 · Xingang Guo, Bin Hu

Value-based methods play a fundamental role in Markov decision processes (MDPs) and reinforcement learning (RL). In this paper, we present a unified control-theoretic framework for analyzing valued-based methods such as …

Reinforcement Learning (RL)

CDiNN -Convex Difference Neural Networks

2021-03-31 · Parameswaran Sankaranarayanan, Raghunathan Rengaswamy

Neural networks with ReLU activation function have been shown to be universal function approximators and learn function mapping as non-smooth functions. Recently, there is considerable interest in the use of neural netwo…

Optimization of stochastic switching buffer network via DC programming

2022-07-18 · Chengyan Zhao, Kazunori Sakurama, Masaki Ogura

This letter deals with the optimization problems of stochastic switching buffer networks, where the switching law is governed by Markov process. The dynamical buffer network is introduced, and its application in modeling…