paper-with-me

Papers

Robust Policy Optimization in Continuous-time Mixed $\mathcal{H}_2/\mathcal{H}_\infty$ Stochastic Control

2022-09-09 · Leilei Cui, Lekan Molu

Following the recent resurgence in establishing linear control theoretic benchmarks for reinforcement leaning (RL)-based policy optimization (PO) for complex dynamical systems with continuous state and action spaces, an optimal control problem for a continuous-time infinite-dimensional linear stochastic system possessing additive Brownian motion is optimized on a cost that is an exponent of the quadratic form of the state, input, and disturbance terms. We lay out a model-based and model-free algorithm for RL-based stochastic PO. For the model-based algorithm, we establish rigorous convergence guarantees. For the sampling-based algorithm, over trajectory arcs that emanate from the phase space, we find that the Hamilton-Jacobi Bellman equation parameterizes trajectory costs -- resulting in a discrete-time (input and state-based) sampling scheme accompanied by unknown nonlinear dynamics with continuous-time policy iterates. The need for known dynamics operators is circumvented and we arrive at a reinforced PO algorithm (via policy iteration) where an upper bound on the $\mathcal{H}_2$ norm is minimized (to guarantee stability) and a robustness metric is enforced by maximizing the cost with respect to a controller that includes the level of noise attenuation specified by the system's $H_\infty$ norm. Rigorous robustness analyses is prescribed in an input-to-state stability formalism. Our analyses and contributions are distinguished by many natural systems characterized by additive Wiener process, amenable to \^Ito's stochastic differential calculus in dynamic game settings.

📄 PDF Abstract BibTeX arXiv:2209.04477

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Entropy annealing for policy mirror descent in continuous time and space

2024-05-30 · Deven Sethi, David Šiška, Yufei Zhang

Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additional regularization bias. This work quan…

Policy Gradient Methods

Global Convergence of Direct Policy Search for State-Feedback $\mathcal{H}_\infty$ Robust Control: A Revisit of Nonsmooth Synthesis with Goldstein Subdifferential

2022-10-20 · Xingang Guo, Bin Hu

Direct policy search has been widely applied in modern reinforcement learning and continuous control. However, the theoretical properties of direct policy search on nonsmooth robust control synthesis have not been fully …

continuous-controlContinuous Control

Policy Gradient for Continuous-Time Robust Markov Decision Processes

2026-06-03 · Tanya Veeravalli, David M. Bossens, Atsushi Nitanda arxiv

The framework of robust Markov decision processes (RMDPs) allows the design of reinforcement learning agents that satisfy performance guarantees under worst-case transition dynamics. Traditional RMDPs consider discrete-t…

Reinforcement Learning

$\mathcal{H}_{\infty}$-optimal Interval Observer Synthesis for Uncertain Nonlinear Dynamical Systems via Mixed-Monotone Decompositions

2022-03-14 · Mohammad Khajenejad, Sze Zheng Yong

This paper introduces a novel $\mathcal{H}_{\infty}$-optimal interval observer synthesis for bounded-error/uncertain locally Lipschitz nonlinear continuous-time (CT) and discrete-time (DT) systems with noisy nonlinear ob…

Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation

2025-05-20 · Runze Zhao, Yue Yu, Adams Yiyue Zhu, Chen Yang 외

Continuous-time reinforcement learning (CTRL) provides a principled framework for sequential decision-making in environments where interactions evolve continuously over time. Despite its empirical success, the theoretica…

Computational Efficiencycontinuous-controlContinuous Controlreinforcement-learning+2