paper-with-me

Papers

Active Reward Learning from Online Preferences

2023-02-27 · Vivek Myers, Erdem Biyik, Dorsa Sadigh

Robot policies need to adapt to human preferences and/or new environments. Human experts may have the domain knowledge required to help robots achieve this adaptation. However, existing works often require costly offline re-training on human feedback, and those feedback usually need to be frequent and too complex for the humans to reliably provide. To avoid placing undue burden on human experts and allow quick adaptation in critical real-world situations, we propose designing and sparingly presenting easy-to-answer pairwise action preference queries in an online fashion. Our approach designs queries and determines when to present them to maximize the expected value derived from the queries' information. We demonstrate our approach with experiments in simulation, human user studies, and real robot experiments. In these settings, our approach outperforms baseline techniques while presenting fewer queries to human experts. Experiment videos, code and appendices are found at https://sites.google.com/view/onlineactivepreferences.

📄 PDF Abstract BibTeX arXiv:2302.13507

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Learning Reward Functions from User Interactions

2017-08-15 · Li Ziming, Kiseleva Julia, de Rijke Maarten, Grotov Artem

In the physical world, people have dynamic preferences, e.g., the same situation can lead to satisfaction for some humans and to frustration for others. Personalization is called for. The same observation holds for onlin…

Recommendation Systems

PAPA: Online Personalized Active Preference Alignment

2026-07-01 · Anindya Sarkar, Nasik Muhammad Nafi, Isaac Lyngaas, Muralikrishnan Gopalakrishnan Meena 외 arxiv

Diffusion models are highly effective at modeling complex data distributions, including images and text. However, in applications like personalized recommender systems, the objective often shifts to modeling specific reg…

Reinforcement Learning

Incorporating Behavioral Constraints in Online AI Systems

2018-09-15 · Avinash Balakrishnan, Djallel Bouneffouf, Nicholas Mattei, Francesca Rossi

AI systems that learn through reward feedback about the actions they take are increasingly deployed in domains that have significant impact on our daily life. However, in many cases the online rewards should not be the o…

Thompson Sampling

Online Policy Learning from Offline Preferences

2024-03-15 · Guoxi Zhang, Han Bao, Hisashi Kashima

In preference-based reinforcement learning (PbRL), a reward function is learned from a type of human feedback called preference. To expedite preference collection, recent works have leveraged \emph{offline preferences}, …

continuous-controlContinuous Control

MORAL: Aligning AI with Human Norms through Multi-Objective Reinforced Active Learning

2021-12-30 · Markus Peschl, Arkady Zgonnikov, Frans A. Oliehoek, Luciano C. Siebert

Inferring reward functions from demonstrations and pairwise preferences are auspicious approaches for aligning Reinforcement Learning (RL) agents with human intentions. However, state-of-the art methods typically focus o…

Active LearningEthicsReinforcement Learning (RL)