paper-with-me

홈 › Papers

Off-Policy Evaluation of Ranking Policies under Diverse User Behavior

2023-06-26 · Haruka Kiyohara, Masatoshi Uehara, Yusuke Narita, Nobuyuki Shimizu, Yasuo Yamamoto, Yuta Saito

Ranking interfaces are everywhere in online platforms. There is thus an ever growing interest in their Off-Policy Evaluation (OPE), aiming towards an accurate performance evaluation of ranking policies using logged data. A de-facto approach for OPE is Inverse Propensity Scoring (IPS), which provides an unbiased and consistent value estimate. However, it becomes extremely inaccurate in the ranking setup due to its high variance under large action spaces. To deal with this problem, previous studies assume either independent or cascade user behavior, resulting in some ranking versions of IPS. While these estimators are somewhat effective in reducing the variance, all existing estimators apply a single universal assumption to every user, causing excessive bias and variance. Therefore, this work explores a far more general formulation where user behavior is diverse and can vary depending on the user context. We show that the resulting estimator, which we call Adaptive IPS (AIPS), can be unbiased under any complex user behavior. Moreover, AIPS achieves the minimum variance among all unbiased estimators based on IPS. We further develop a procedure to identify the appropriate user behavior model to minimize the mean squared error (MSE) of AIPS in a data-driven fashion. Extensive experiments demonstrate that the empirical accuracy improvement can be significant, enabling effective OPE of ranking systems even under diverse user behavior.

📄 PDF Abstract BibTeX arXiv:2306.15098

Code (1)

aiueola/kdd2023-aips 공식 구현

Tasks

Off-policy evaluation

Similar Papers 제목 키워드 기반

Ranking Policy Decisions

2020-08-31 · NeurIPS 2021 12 · Hadrien Pouget, Hana Chockler, Youcheng Sun, Daniel Kroening

Policies trained via Reinforcement Learning (RL) are often needlessly complex, making them difficult to analyse and interpret. In a run with $n$ time steps, a policy will make $n$ decisions on actions to take; we conject…

Atari GamesReinforcement Learning (RL)

Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies

2026-03-23 · Koichi Tanaka, Kazuki Kawamura, Takanori Muroi, Yusuke Narita 외 arxiv

Off-Policy Evaluation (OPE) is an important practical problem in algorithmic ranking systems, where the goal is to estimate the expected performance of a new ranking policy using only offline logged data collected under …

Offline Policy Selection under Uncertainty

2020-12-12 · Mengjiao Yang, Bo Dai, Ofir Nachum, George Tucker 외

The presence of uncertainty in policy evaluation significantly complicates the process of policy ranking and selection in real-world settings. We formally consider offline policy selection as learning preferences over a …

Validate on Sim, Detect on Real -- Model Selection for Domain Randomization

2021-11-01 · Gal Leibovich, Guy Jacob, Shadi Endrawis, Gal Novik 외

A practical approach to learning robot skills, often termed sim2real, is to train control policies in simulation and then deploy them on a real robot. Popular techniques to improve the sim2real transfer build on domain r…

Model SelectionOut-of-Distribution DetectionRobotic Grasping

Supervised Off-Policy Ranking

2021-07-03 · Yue Jin, Yue Zhang, Tao Qin, Xudong Zhang 외

Off-policy evaluation (OPE) is to evaluate a target policy with data generated by other policies. Most previous OPE methods focus on precisely estimating the true performance of a policy. We observe that in many applicat…

Off-policy evaluation