paper-with-me

홈 › Papers

Policy Optimization for Continuous Reinforcement Learning

2023-05-30 · NeurIPS 2023 11

We study reinforcement learning (RL) in the setting of continuous time and space, for an infinite horizon with a discounted objective and the underlying dynamics driven by a stochastic differential equation. Built upon recent advances in the continuous approach to RL, we develop a notion of occupation time (specifically for a discounted objective), and show how it can be effectively used to derive performance-difference and local-approximation formulas. We further extend these results to illustrate their applications in the PG (policy gradient) and TRPO/PPO (trust region policy optimization/ proximal policy optimization) methods, which have been familiar and powerful tools in the discrete RL setting but under-developed in continuous RL. Through numerical experiments, we demonstrate the effectiveness and advantages of our approach.

📄 PDF Abstract BibTeX arXiv:2305.18901

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Quasi-Newton Trust Region Policy Optimization

2019-12-26 · Devesh Jha, Arvind Raghunathan, Diego Romeres

We propose a trust region method for policy optimization that employs Quasi-Newton approximation for the Hessian, called Quasi-Newton Trust Region Policy Optimization QNTRPO. Gradient descent is the de facto algorithm fo…

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning+1

A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models

2026-07-16 · Zikun Zhang, Jiayuan Sheng, David D. Yao, Wenpin Tang arxiv

We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuo…

Mathematical ReasoningReinforcement Learning

Deep Reinforcement Learning for Stock Portfolio Optimization

2020-12-09 · Le Trung Hieu

Stock portfolio optimization is the process of constant re-distribution of money to a pool of various stocks. In this paper, we will formulate the problem such that we can apply Reinforcement Learning for the task proper…

Deep Reinforcement LearningPortfolio Optimizationreinforcement-learningReinforcement Learning+1

Iterative Amortized Policy Optimization

2020-10-20 · NeurIPS 2021 12 · Joseph Marino, Alexandre Piché, Alessandro Davide Ialongo, Yisong Yue

Policy networks are a central feature of deep reinforcement learning (RL) algorithms for continuous control, enabling the estimation and sampling of high-value actions. From the variational inference perspective on RL, p…

continuous-controlContinuous ControlDeep Reinforcement Learningreinforcement-learning+2

Comparing Deep Reinforcement Learning and Evolutionary Methods in Continuous Control

2017-11-30 · Shangtong Zhang, Osmar R. Zaiane

Reinforcement Learning and the Evolutionary Strategy are two major approaches in addressing complicated control problems. Both are strong contenders and have their own devotee communities. Both groups have been very acti…

continuous-controlContinuous ControlDeep Reinforcement Learningreinforcement-learning+2