paper-with-me

홈 › Papers

Estimating Error and Bias in Offline Evaluation Results

2020-01-26 · Mucun Tian, Michael D. Ekstrand

Offline evaluations of recommender systems attempt to estimate users' satisfaction with recommendations using static data from prior user interactions. These evaluations provide researchers and developers with first approximations of the likely performance of a new system and help weed out bad ideas before presenting them to users. However, offline evaluation cannot accurately assess novel, relevant recommendations, because the most novel items were previously unknown to the user, so they are missing from the historical data and cannot be judged as relevant. We present a simulation study to estimate the error that such missing data causes in commonly-used evaluation metrics in order to assess its prevalence and impact. We find that missing data in the rating or observation process causes the evaluation protocol to systematically mis-estimate metric values, and in some cases erroneously determine that a popularity-based recommender outperforms even a perfect personalized recommender. Substantial breakthroughs in recommendation quality, therefore, will be difficult to assess with existing offline techniques.

📄 PDF Abstract BibTeX arXiv:2001.09455

Code (1)

BoiseState/chiir2020-est-error 공식 구현

Tasks

Recommendation Systems

Similar Papers 제목 키워드 기반

A More Accurate Algorithm Comparison through A/B Testing using Offline Evaluation Methods

2026-07-02 · Koki Konishi, Masataka Ushiku, Yuta Saito arxiv

A/B testing is the gold standard for selecting the better algorithm in online services. While offline evaluation has attracted attention as a safer alternative due to the high experimental costs and the potential risk of…

The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice

2026-05-02 · Jikai Jin, Vasilis Syrgkanis arxiv

Offline evaluation of language models from usage logs is biased when model choice is confounded: the same user-side factors that influence which model is used can also influence how its output is judged, so raw compariso…

Distributional Offline Policy Evaluation with Predictive Error Guarantees

2023-02-19 · Runzhe Wu, Masatoshi Uehara, Wen Sun

We study the problem of estimating the distribution of the return of a policy using an offline dataset that is not generated from the policy, i.e., distributional offline policy evaluation (OPE). We propose an algorithm …

On Instrumental Variable Regression for Deep Offline Policy Evaluation

2021-05-21 · Yutian Chen, Liyuan Xu, Caglar Gulcehre, Tom Le Paine 외

We show that the popular reinforcement learning (RL) strategy of estimating the state-action value (Q-function) by minimizing the mean squared Bellman error leads to a regression problem with confounding, the inputs and …

regressionReinforcement Learning (RL)

Exclusively Penalized Q-learning for Offline Reinforcement Learning

2024-05-23 · Junghyuk Yeom, Yonghyeon Jo, Jungmo Kim, Sanghyeon Lee 외

Constraint-based offline reinforcement learning (RL) involves policy constraints or imposing penalties on the value function to mitigate overestimation errors caused by distributional shift. This paper focuses on a limit…

Offline RLQ-Learningreinforcement-learningReinforcement Learning+1