An Online Model-Following Projection Mechanism Using Reinforcement Learning
In this paper, we propose a model-free adaptive learning solution for a model-following control problem. This approach employs policy iteration, to find an optimal adaptive control solution. It utilizes a moving finite-horizon of model-following error measurements. In addition, the control strategy is designed by using a projection mechanism that employs Lagrange dynamics. It allows for real-time tuning of derived actor-critic structures to find the optimal model-following strategy and sustain optimized adaptation performance. Finally, the efficacy of the proposed framework is emphasized through a comparison with sliding mode and high-order model-free adaptive control approaches. Keywords: Model Reference Adaptive Systems, Reinforcement Learning, adaptive critics, control system, stochastic, nonlinear system
Code (0)
등록된 구현이 없습니다.
Tasks
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
An Observer-Based Reinforcement Learning Solution for Model-Following Problems
In this paper, a multi-objective model-following control problem is solved using an observer-based adaptive learning scheme. The overall goal is to regulate the model-following error dynamics along with optimizing the dy…
reinforcement-learningReinforcement LearningProvably Safe Reinforcement Learning via Action Projection using Reachability Analysis and Polynomial Zonotopes
While reinforcement learning produces very promising results for many applications, its main disadvantage is the lack of safety guarantees, which prevents its use in safety-critical systems. In this work, we address this…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement LearningThe Real, the Better: Aligning Large Language Models with Online Human Behaviors
Large language model alignment is widely used and studied to avoid LLM producing unhelpful and harmful responses. However, the lengthy training process and predefined preference bias hinder adaptation to online diverse h…
Language ModelingLanguage ModellingLarge Language ModelFrom Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning
Safety in reinforcement learning is often specified through cumulative cost constraints, but these trajectory-level guarantees do not directly prevent unsafe individual decisions, especially under nonstationarity. In con…
Reinforcement LearningVisualizing Critic Match Loss Landscapes for Interpretation of Online Reinforcement Learning Control Algorithms
Reinforcement learning has proven its power on various occasions. However, its performance is not always guaranteed when system dynamics change. Instead, it largely relies on users' empirical experience. For reinforcemen…
Reinforcement Learning