paper-with-me

홈 › Papers

On the Value of Bandit Feedback for Offline Recommender System Evaluation

2019-07-26 · Olivier Jeunen, David Rohde, Flavian vasile

In academic literature, recommender systems are often evaluated on the task of next-item prediction. The procedure aims to give an answer to the question: "Given the natural sequence of user-item interactions up to time t, can we predict which item the user will interact with at time t+1?". Evaluation results obtained through said methodology are then used as a proxy to predict which system will perform better in an online setting. The online setting, however, poses a subtly different question: "Given the natural sequence of user-item interactions up to time t, can we get the user to interact with a recommended item at time t+1?". From a causal perspective, the system performs an intervention, and we want to measure its effect. Next-item prediction is often used as a fall-back objective when information about interventions and their effects (shown recommendations and whether they received a click) is unavailable. When this type of data is available, however, it can provide great value for reliably estimating online recommender system performance. Through a series of simulated experiments with the RecoGym environment, we show where traditional offline evaluation schemes fall short. Additionally, we show how so-called bandit feedback can be exploited for effective offline evaluation that more accurately reflects online performance.

📄 PDF Abstract BibTeX arXiv:1907.12384

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation Systems

Similar Papers 제목 키워드 기반

BanditMF: Multi-Armed Bandit Based Matrix Factorization Recommender System

2021-06-21 · Shenghao Xu

Multi-armed bandits (MAB) provide a principled online learning approach to attain the balance between exploration and exploitation. Due to the superior performance and low feedback learning without the learning to act in…

Collaborative FilteringMulti-Armed BanditsRecommendation Systemsvalid

Online Matching: A Real-time Bandit System for Large-scale Recommendations

2023-07-29 · Xinyang Yi, Shao-Chuan Wang, Ruining He, Hariharan Chandrasekaran 외

The last decade has witnessed many successes of deep learning-based models for industry-scale recommender systems. These models are typically trained offline in a batch manner. While being effective in capturing users' p…

Multi-Armed BanditsRecommendation Systems

Comparison-based Conversational Recommender System with Relative Bandit Feedback

2022-08-21 · Zhihui Xie, Tong Yu, Canzhe Zhao, Shuai Li

With the recent advances of conversational recommendations, the recommender system is able to actively and dynamically elicit user preference via conversational interactions. To achieve this, the system periodically quer…

Recommendation Systems

Deep Bayesian Bandits: Exploring in Online Personalized Recommendations

2020-08-03 · Dalin Guo, Sofia Ira Ktena, Ferenc Huszar, Pranay Kumar Myana 외

Recommender systems trained in a continuous learning fashion are plagued by the feedback loop problem, also known as algorithmic bias. This causes a newly trained model to act greedily and favor items that have already b…

Recommendation Systems

Existence conditions for hidden feedback loops in online recommender systems

2021-09-11 · Anton S. Khritankov, Anton A. Pilkevich

We explore a hidden feedback loops effect in online recommender systems. Feedback loops result in degradation of online multi-armed bandit (MAB) recommendations to a small subset and loss of coverage and novelty. We stud…

Recommendation Systems