paper-with-me

홈 › Papers

Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback

2024-01-17 · Teng Xiao, Suhang Wang

Probabilistic learning to rank (LTR) has been the dominating approach for optimizing the ranking metric, but cannot maximize long-term rewards. Reinforcement learning models have been proposed to maximize user long-term rewards by formulating the recommendation as a sequential decision-making problem, but could only achieve inferior accuracy compared to LTR counterparts, primarily due to the lack of online interactions and the characteristics of ranking. In this paper, we propose a new off-policy value ranking (VR) algorithm that can simultaneously maximize user long-term rewards and optimize the ranking metric offline for improved sample efficiency in a unified Expectation-Maximization (EM) framework. We theoretically and empirically show that the EM process guides the leaned policy to enjoy the benefit of integration of the future reward and ranking metric, and learn without any online interactions. Extensive offline and online experiments demonstrate the effectiveness of our methods.

📄 PDF Abstract BibTeX arXiv:2401.08959

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingLearning-To-Rankreinforcement-learningSequential Decision Making

Similar Papers 제목 키워드 기반

Navigating Noisy Feedback: Enhancing Reinforcement Learning with Error-Prone Language Models

2024-10-22 · MuHan Lin, Shuyang Shi, Yue Guo, Behdad Chalaki 외

The correct specification of reward models is a well-known challenge in reinforcement learning. Hand-crafted reward functions often lead to inefficient or suboptimal policies and may not be aligned with user values. Rein…

HallucinationLanguage ModelingLanguage ModellingLarge Language Model+2

Zeroth-Order Optimization Meets Human Feedback: Provable Learning via Ranking Oracles

2023-03-07 · Zhiwei Tang, Dmitry Rybin, Tsung-Hui Chang

In this study, we delve into an emerging optimization challenge involving a black-box objective function that can only be gauged via a ranking oracle-a situation frequently encountered in real-world scenarios, especially…

Image Generationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Policy-Gradient Training of Fair and Unbiased Ranking Functions

2019-11-19 · Himank Yadav, Zhengxiao Du, Thorsten Joachims

While implicit feedback (e.g., clicks, dwell times, etc.) is an abundant and attractive source of data for learning to rank, it can produce unfair ranking policies for both exogenous and endogenous reasons. Exogenous rea…

counterfactualDecision MakingFairnessLearning-To-Rank

APRIL: Active Preference-learning based Reinforcement Learning

2012-08-05 · Riad Akrour, Marc Schoenauer, Michèle Sebag

This paper focuses on reinforcement learning (RL) with limited prior knowledge. In the domain of swarm robotics for instance, the expert can hardly design a reward function or demonstrate the target behavior, forbidding …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Nash Learning from Human Feedback

2023-12-01 · Rémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar 외

Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences. Typically, RLHF involves the initial step of learning a reward model fr…

Text Summarization