Learned Ranking Function: From Short-term Behavior Predictions to Long-term User Satisfaction
We present the Learned Ranking Function (LRF), a system that takes short-term user-item behavior predictions as input and outputs a slate of recommendations that directly optimizes for long-term user satisfaction. Most previous work is based on optimizing the hyperparameters of a heuristic function. We propose to model the problem directly as a slate optimization problem with the objective of maximizing long-term user satisfaction. We also develop a novel constraint optimization algorithm that stabilizes objective trade-offs for multi-objective optimization. We evaluate our approach with live experiments and describe its deployment on YouTube.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
On short-time behavior of implied volatility in a market model with indexes
This paper investigates short-term behaviors of implied volatility of derivatives written on indexes in equity markets when the index processes are constructed by using a ranking procedure. Even in simple market settings…
On the Value of Myopic Behavior in Policy Reuse
Leveraging learned strategies in unfamiliar scenarios is fundamental to human intelligence. In reinforcement learning, rationally reusing the policies acquired from other tasks or human experts is critical for tackling p…
Multi-Level Interaction Reranking with User Behavior History
As the final stage of the multi-stage recommender system (MRS), reranking directly affects users' experience and satisfaction, thus playing a critical role in MRS. Despite the improvement achieved in the existing work, t…
Recommendation SystemsRerankingTowards Axiomatic Explanations for Neural Ranking Models
Recently, neural networks have been successfully employed to improve upon state-of-the-art performance in ad-hoc retrieval tasks via machine-learned ranking functions. While neural retrieval models grow in complexity and…
Document RankingInformation RetrievalRetrievalLearning diverse rankings with multi-armed bandits
Algorithms for learning to rank Web documents usually assume a document's relevance is independent of other documents. This leads to learned ranking functions that produce rankings with redundant results. In contrast, us…
DiversityLearning-To-RankMulti-Armed Bandits