paper-with-me

홈 › Papers

From self-tuning regulators to reinforcement learning and back again

2019-06-27 · Nikolai Matni, Alexandre Proutiere, Anders Rantzer, Stephen Tu

Machine and reinforcement learning (RL) are increasingly being applied to plan and control the behavior of autonomous systems interacting with the physical world. Examples include self-driving vehicles, distributed sensor networks, and agile robots. However, when machine learning is to be applied in these new settings, the algorithms had better come with the same type of reliability, robustness, and safety bounds that are hallmarks of control theory, or failures could be catastrophic. Thus, as learning algorithms are increasingly and more aggressively deployed in safety critical settings, it is imperative that control theorists join the conversation. The goal of this tutorial paper is to provide a starting point for control theorists wishing to work on learning related problems, by covering recent advances bridging learning and control theory, and by placing these results within an appropriate historical context of system identification and adaptive control.

📄 PDF Abstract BibTeX arXiv:1906.11392

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Advanced control based on Recurrent Neural Networks learned using Virtual Reference Feedback Tuning and application to an Electronic Throttle Body (with supplementary material)

2021-03-03 · William D'Amico, Marcello Farina, Giulio Panzani

In this paper the application of Virtual Reference Feedback Tuning (VRFT) for control of nonlinear systems with regulators defined by Echo State Networks (ESN) and Long Short Term Memory (LSTM) networks is investigated. …

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals

2025-02-14 · Andrew Kiruluta, Andreas Lemos, Priscilla Burity

We propose a novel reinforcement learning framework for post training large language models that does not rely on human in the loop feedback. Instead, our approach uses cross attention signals within the model itself to …

Policy Gradient Methods

On-Policy Robust Adaptive Discrete-Time Regulator for Passive Unidirectional System using Stochastic Hill-climbing Algorithm and Associated Search Element

2021-12-30 · Mohsen Jafarzadeh, Nicholas Gans, Yonas Tadesse

Non-linear discrete-time state-feedback regulators are widely used in passive unidirectional systems. Offline system identification is required for tuning parameters of these regulators. However, offline system identific…

Reinforcement Learning Fine-tuning of Language Models is Biased Towards More Extractable Features

2023-11-07 · Diogo Cruz, Edoardo Pona, Alex Holness-Tofts, Elias Schmied 외

Many capable large language models (LLMs) are developed via self-supervised pre-training followed by a reinforcement-learning fine-tuning phase, often based on human or AI feedback. During this stage, models may be guide…

reinforcement-learningReinforcement Learning

Online Learning from Strategic Human Feedback in LLM Fine-Tuning

2024-12-22 · Shugang Hao, Lingjie Duan

Reinforcement learning from human feedback (RLHF) has become an essential step in fine-tuning large language models (LLMs) to align them with human preferences. However, human labelers are selfish and have diverse prefer…