Reward Shaping for User Satisfaction in a REINFORCE Recommender
How might we design Reinforcement Learning (RL)-based recommenders that encourage aligning user trajectories with the underlying user satisfaction? Three research questions are key: (1) measuring user satisfaction, (2) combatting sparsity of satisfaction signals, and (3) adapting the training of the recommender agent to maximize satisfaction. For measurement, it has been found that surveys explicitly asking users to rate their experience with consumed items can provide valuable orthogonal information to the engagement/interaction data, acting as a proxy to the underlying user satisfaction. For sparsity, i.e, only being able to observe how satisfied users are with a tiny fraction of user-item interactions, imputation models can be useful in predicting satisfaction level for all items users have consumed. For learning satisfying recommender policies, we postulate that reward shaping in RL recommender agents is powerful for driving satisfying user experiences. Putting everything together, we propose to jointly learn a policy network and a satisfaction imputation network: The role of the imputation network is to learn which actions are satisfying to the user; while the policy network, built on top of REINFORCE, decides which items to recommend, with the reward utilizing the imputed satisfaction. We use both offline analysis and live experiments in an industrial large-scale recommendation platform to demonstrate the promise of our approach for satisfying user experiences.
Code (0)
등록된 구현이 없습니다.
Tasks
ImputationReinforcement Learning (RL)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
ROLeR: Effective Reward Shaping in Offline Reinforcement Learning for Recommender Systems
Offline reinforcement learning (RL) is an effective tool for real-world recommender systems with its capacity to model the dynamic interest of users and its interactive nature. Most existing offline RL recommender system…
Offline RLRecommendation Systemsreinforcement-learningReinforcement Learning+1Explicit User Manipulation in Reinforcement Learning Based Recommender Systems
Recommender systems are highly prevalent in the modern world due to their value to both users and platforms and services that employ them. Generally, they can improve the user experience and help to increase satisfaction…
Recommendation Systemsreinforcement-learningReinforcement LearningReinforcement Learning (RL)Reinforcement Learning-based Recommender Systems with Large Language Models for State Reward and Action Modeling
Reinforcement Learning (RL)-based recommender systems have demonstrated promising performance in meeting user expectations by learning to make accurate next-item recommendations from historical user-item interactions. Ho…
Offline RLRecommendation SystemsReinforcement Learning (RL)Sequential RecommendationAdversarial Batch Inverse Reinforcement Learning: Learn to Reward from Imperfect Demonstration for Interactive Recommendation
Rewards serve as a measure of user satisfaction and act as a limiting factor in interactive recommender systems. In this research, we focus on the problem of learning to reward (LTR), which is fundamental to reinforcemen…
Interactive RecommendationRecommendation Systemsreinforcement-learningReinforcement LearningMulti-Task Fusion via Reinforcement Learning for Long-Term User Satisfaction in Recommender Systems
Recommender System (RS) is an important online application that affects billions of users every day. The mainstream RS ranking framework is composed of two parts: a Multi-Task Learning model (MTL) that predicts various u…
Multi-Task LearningRecommendation SystemsReinforcement Learning (RL)