paper-with-me

Papers

Control randomisation approach for policy gradient and application to reinforcement learning in optimal switching

2024-04-27 · Robert Denkert, Huyên Pham, Xavier Warin

We propose a comprehensive framework for policy gradient methods tailored to continuous time reinforcement learning. This is based on the connection between stochastic control problems and randomised problems, enabling applications across various classes of Markovian continuous time control problems, beyond diffusion models, including e.g. regular, impulse and optimal stopping/switching problems. By utilizing change of measure in the control randomisation technique, we derive a new policy gradient representation for these randomised problems, featuring parametrised intensity policies. We further develop actor-critic algorithms specifically designed to address general Markovian stochastic control issues. Our framework is demonstrated through its application to optimal switching problems, with two numerical case studies in the energy sector focusing on real options.

📄 PDF Abstract BibTeX arXiv:2404.17939

Code (0)

등록된 구현이 없습니다.

Tasks

Policy Gradient Methods

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Robust Domain Randomised Reinforcement Learning through Peer-to-Peer Distillation

2020-12-09 · Chenyang Zhao, Timothy Hospedales

In reinforcement learning, domain randomisation is an increasingly popular technique for learning more general policies that are robust to domain-shifts at deployment. However, naively aggregating information from random…

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning+1

Deep combinatorial optimisation for optimal stopping time problems : application to swing options pricing

2020-01-30 · Thomas Deschatre, Joseph Mikael

A new method for stochastic control based on neural networks and using randomisation of discrete random variables is proposed and applied to optimal stopping time problems. The method models directly the policy and does …

Solving the Real Robot Challenge using Deep Reinforcement Learning

2021-09-30 · Robert McCarthy, Francisco Roldan Sanchez, Qiang Wang, David Cordova Bulens 외

This paper details our winning submission to Phase 1 of the 2021 Real Robot Challenge; a challenge in which a three-fingered robot must carry a cube along specified goal trajectories. To solve Phase 1, we use a pure rein…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

Analysing Deep Reinforcement Learning Agents Trained with Domain Randomisation

2019-12-18 · Tianhong Dai, Kai Arulkumaran, Tamara Gerbert, Samyakh Tukra 외

Deep reinforcement learning has the potential to train robots to perform complex tasks in the real world without requiring accurate models of the robot or its environment. A practical approach is to train agents in simul…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Towards a Theoretical Foundation of Policy Optimization for Learning Control Policies

2022-10-10 · Bin Hu, Kaiqing Zhang, Na Li, Mehran Mesbahi 외

Gradient-based methods have been widely used for system design and optimization in diverse application domains. Recently, there has been a renewed interest in studying theoretical properties of these methods in the conte…

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning+1