paper-with-me

홈 › Papers

Counterfactual Evaluation of Slate Recommendations with Sequential Reward Interactions

2020-07-25 · James McInerney, Brian Brost, Praveen Chandar, Rishabh Mehrotra, Ben Carterette

Users of music streaming, video streaming, news recommendation, and e-commerce services often engage with content in a sequential manner. Providing and evaluating good sequences of recommendations is therefore a central problem for these services. Prior reweighting-based counterfactual evaluation methods either suffer from high variance or make strong independence assumptions about rewards. We propose a new counterfactual estimator that allows for sequential interactions in the rewards with lower variance in an asymptotically unbiased manner. Our method uses graphical assumptions about the causal relationships of the slate to reweight the rewards in the logging policy in a way that approximates the expected sum of rewards under the target policy. Extensive experiments in simulation and on a live recommender system show that our approach outperforms existing methods in terms of bias and data efficiency for the sequential track recommendations problem.

📄 PDF Abstract BibTeX arXiv:2007.12986

Code (1)

spotify-research/RIPS_KDD2020 공식 구현

Tasks

counterfactualNews RecommendationOff-policy evaluationRecommendation Systems

Similar Papers 제목 키워드 기반

Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation

2022-08-10 · Imad Aouali, Achraf Ait Sidi Hammou, Otmane Sakhi, David Rohde 외

We introduce Probabilistic Rank and Reward (PRR), a scalable probabilistic model for personalized slate recommendation. Our approach allows off-policy estimation of the reward in the scenario where the user interacts wit…

Recommendation Systems

SlateFree: a Model-Free Decomposition for Reinforcement Learning with Slate Actions

2022-09-05 · Anastasios Giovanidis

We consider the problem of sequential recommendations, where at each step an agent proposes some slate of $N$ distinct items to a user from a much larger catalog of size $K>>N$. The user has unknown preferences towards t…

Q-Learningreinforcement-learningReinforcement Learning (RL)

Improving Sequential Recommenders through Counterfactual Augmentation of System Exposure

2025-04-18 · Ziqi Zhao, Zhaochun Ren, Jiyuan Yang, Zuming Yan 외

In sequential recommendation (SR), system exposure refers to items that are exposed to the user. Typically, only a few of the exposed items would be interacted with by the user. Although SR has achieved great success in …

counterfactualSequential Recommendation

Verifiable Counterfactual Supervision for Process Reward Models

2026-05-04 · Yinghui Chi, Yuanhong Wang arxiv

Process reward models (PRMs) require supervision that identifies not only whether a reasoning trajectory is correct, but also where the reasoning process first becomes unsupported by its prefix. We frame this requirement…

Logical Reasoning

Sequential Counterfactual Decision-Making Under Confounded Reward

2022-06-05 · Erik Skalnes

We investigate the limitations of random trials when the cause of interest is confounded with the effect by formalizing a counterfactual policy-space where the agent's natural predilection is input to a soft-intervention…

counterfactualDecision Making