paper-with-me

Papers

A Policy Gradient Framework for Stochastic Optimal Control Problems with Global Convergence Guarantee

2023-02-11 · Mo Zhou, Jianfeng Lu

We consider policy gradient methods for stochastic optimal control problem in continuous time. In particular, we analyze the gradient flow for the control, viewed as a continuous time limit of the policy gradient method. We prove the global convergence of the gradient flow and establish a convergence rate under some regularity assumptions. The main novelty in the analysis is the notion of local optimal control function, which is introduced to characterize the local optimality of the iterate.

📄 PDF Abstract BibTeX arXiv:2302.05816

Code (0)

등록된 구현이 없습니다.

Tasks

Policy Gradient Methods

Similar Papers 제목 키워드 기반

Control randomisation approach for policy gradient and application to reinforcement learning in optimal switching

2024-04-27 · Robert Denkert, Huyên Pham, Xavier Warin

We propose a comprehensive framework for policy gradient methods tailored to continuous time reinforcement learning. This is based on the connection between stochastic control problems and randomised problems, enabling a…

Policy Gradient Methods

Inferring the Optimal Policy using Markov Chain Monte Carlo

2019-11-16 · Brandon Trabucco, Albert Qu, Simon Li, Ganeshkumar Ashokavardhanan

This paper investigates methods for estimating the optimal stochastic control policy for a Markov Decision Process with unknown transition dynamics and an unknown reward function. This form of model-free reinforcement le…

Reinforcement Learning

Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action

2024-09-25 · Xin Chen, Yifan Hu, Minda Zhao

Policy gradient methods are widely used in reinforcement learning. Yet, the nonconvexity of policy optimization imposes significant challenges in understanding the global convergence of policy gradient methods. For a cla…

Policy Gradient Methods

Policy gradient learning methods for stochastic control with exit time and applications to share repurchase pricing

2023-02-14 · Mohamed Hamdouche, Pierre Henry-Labordere, Huyen Pham

We develop policy gradients methods for stochastic control with exit time in a model-free setting. We propose two types of algorithms for learning either directly the optimal policy or by learning alternately the value f…

Policy Gradient Methods

Neural Policy Iteration for Stochastic Optimal Control: A Physics-Informed Approach

2025-08-03 · Yeongjong Kim, Yeoneung Kim, Minseok Kim, Namkyeong Cho arxiv

We propose a physics-informed neural network policy iteration (PINN-PI) framework for solving stochastic optimal control problems governed by second-order Hamilton--Jacobi--Bellman (HJB) equations. At each iteration, a n…