paper-with-me

홈 › Papers

Dueling RL: Reinforcement Learning with Trajectory Preferences

2021-11-08 · Aldo Pacchiano, Aadirupa Saha, Jonathan Lee

We consider the problem of preference based reinforcement learning (PbRL), where, unlike traditional reinforcement learning, an agent receives feedback only in terms of a 1 bit (0/1) preference over a trajectory pair instead of absolute rewards for them. The success of the traditional RL framework crucially relies on the underlying agent-reward model, which, however, depends on how accurately a system designer can express an appropriate reward function and often a non-trivial task. The main novelty of our framework is the ability to learn from preference-based trajectory feedback that eliminates the need to hand-craft numeric reward models. This paper sets up a formal framework for the PbRL problem with non-markovian rewards, where the trajectory preferences are encoded by a generalized linear model of dimension $d$. Assuming the transition model is known, we then propose an algorithm with almost optimal regret guarantee of $\tilde {\mathcal{O}}\left( SH d \log (T / \delta) \sqrt{T} \right)$. We further, extend the above algorithm to the case of unknown transition dynamics, and provide an algorithm with near optimal regret guarantee $\widetilde{\mathcal{O}}((\sqrt{d} + H^2 + |\mathcal{S}|)\sqrt{dT} +\sqrt{|\mathcal{S}||\mathcal{A}|TH} )$. To the best of our knowledge, our work is one of the first to give tight regret guarantees for preference based RL problems with trajectory preferences.

📄 PDF Abstract BibTeX arXiv:2111.04850

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Breaking the Cold-Start Barrier: Reinforcement Learning with Double and Dueling DQNs

2025-08-28 · Minda Zhao arxiv

Recommender systems struggle to provide accurate suggestions to new users with limited interaction history, a challenge known as the cold-user problem. This paper proposes a reinforcement learning approach using Double a…

Reinforcement LearningActive Learning

Versatile Dueling Bandits: Best-of-both-World Analyses for Online Learning from Preferences

2022-02-14 · Aadirupa Saha, Pierre Gaillard

We study the problem of $K$-armed dueling bandit for both stochastic and adversarial environments, where the goal of the learner is to aggregate information through relative preferences of pair of decisions points querie…

Multi-Armed Bandits

A Minimaximalist Approach to Reinforcement Learning from Human Feedback

2024-01-08 · Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu 외

We present Self-Play Preference Optimization (SPO), an algorithm for reinforcement learning from human feedback. Our approach is minimalist in that it does not require training a reward model nor unstable adversarial tra…

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning

A State Representation Dueling Network for Deep Reinforcement Learning

2020-12-24 · Haomin Qiu, Feng Liu

In recent years there have been many successes in boosting the performance of Deep Q-Networks (DQN). Dueling DQN uses simple dueling architecture but significantly improves the performance of DQN [1]. However, Dueling DQ…

Deep Reinforcement LearningGeneral Reinforcement Learningreinforcement-learningReinforcement Learning+1

Dueling Posterior Sampling for Preference-Based Reinforcement Learning

2019-08-04 · Ellen R. Novoseller, Yibing Wei, Yanan Sui, Yisong Yue 외

In preference-based reinforcement learning (RL), an agent interacts with the environment while receiving preferences instead of absolute feedback. While there is increasing research activity in preference-based RL, the d…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)