paper-with-me

홈 › Papers

Exploiting Unlabeled Data for Feedback Efficient Human Preference based Reinforcement Learning

2023-02-17 · Mudit Verma, Siddhant Bhambri, Subbarao Kambhampati

Preference Based Reinforcement Learning has shown much promise for utilizing human binary feedback on queried trajectory pairs to recover the underlying reward model of the Human in the Loop (HiL). While works have attempted to better utilize the queries made to the human, in this work we make two observations about the unlabeled trajectories collected by the agent and propose two corresponding loss functions that ensure participation of unlabeled trajectories in the reward learning process, and structure the embedding space of the reward model such that it reflects the structure of state space with respect to action distances. We validate the proposed method on one locomotion domain and one robotic manipulation task and compare with the state-of-the-art baseline PEBBLE. We further present an ablation of the proposed loss components across both the domains and find that not only each of the loss components perform better than the baseline, but the synergic combination of the two has much better reward recovery and human feedback sample efficiency.

📄 PDF Abstract BibTeX arXiv:2302.08738

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

2022-03-18 · ICLR 2022 4 · Jongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee 외

Preference-based reinforcement learning (RL) has shown potential for teaching agents to perform the target tasks without a costly, pre-defined reward function by learning the reward with a supervisor's preference between…

Data AugmentationReinforcement Learning (RL)

From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement Learning

2026-05-31 · Jun-Jie Yang, Chia-Heng Hsu, Kui-Yuan Chen, Ping-Chun Hsieh arxiv

Preference-based reinforcement learning (PbRL) avoids explicit reward engineering by learning from pairwise human preference feedback. Existing offline PbRL methods typically follow a two-stage pipeline, first learning a…

Representation LearningReinforcement LearningOffline RL

Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement Learning

2026-06-15 · Tung M. Luu, Hwanhee Kim, Younghwan Lee, Chang D. Yoo arxiv

Conveying complex objectives to reinforcement learning (RL) agents often requires meticulous reward engineering. Preference-based RL (PbRL) offers a promising alternative by learning reward functions from human feedback,…

Reinforcement Learning

COPR: Continual Learning Human Preference through Optimal Policy Regularization

2023-10-24 · Han Zhang, Lin Gui, Yuanzhao Zhai, Hui Wang 외

The technique of Reinforcement Learning from Human Feedback (RLHF) is a commonly employed method to improve pre-trained Language Models (LM), enhancing their ability to conform to human preferences. Nevertheless, the cur…

Continual Learningreinforcement-learningReinforcement Learning

Extensive Self-Contrast Enables Feedback-Free Language Model Alignment

2024-03-31 · Xiao Liu, Xixuan Song, Yuxiao Dong, Jie Tang

Reinforcement learning from human feedback (RLHF) has been a central technique for recent large language model (LLM) alignment. However, its heavy dependence on costly human or LLM-as-Judge preference feedback could stym…

Language ModelingLanguage ModellingLarge Language Modeltext similarity