Counterfactual Learning with General Data-generating Policies
Off-policy evaluation (OPE) attempts to predict the performance of counterfactual policies using log data from a different policy. We extend its applicability by developing an OPE method for a class of both full support and deficient support logging policies in contextual-bandit settings. This class includes deterministic bandit (such as Upper Confidence Bound) as well as deterministic decision-making based on supervised and unsupervised learning. We prove that our method's prediction converges in probability to the true performance of a counterfactual policy as the sample size increases. We validate our method with experiments on partly and entirely deterministic logging policies. Finally, we apply it to evaluate coupon targeting policies by a major online platform and show how to improve the existing policy.
Code (0)
등록된 구현이 없습니다.
Tasks
counterfactualDecision MakingOff-policy evaluationSimilar Papers 제목 키워드 기반
Counterfactual Explanation Policies in RL
As Reinforcement Learning (RL) agents are increasingly employed in diverse decision-making problems using reward preferences, it becomes important to ensure that policies learned by these frameworks in mapping observatio…
counterfactualCounterfactual ExplanationDecision MakingReinforcement Learning (RL)Eventual Discounting Temporal Logic Counterfactual Experience Replay
Linear temporal logic (LTL) offers a simplified way of specifying tasks for policy optimization that may otherwise be difficult to describe with scalar reward functions. However, the standard RL framework can be too myop…
counterfactualCounterfactual ReasoningCounterfactual Explanations for Continuous Action Reinforcement Learning
Reinforcement Learning (RL) has shown great promise in domains like healthcare and robotics but often struggles with adoption due to its lack of interpretability. Counterfactual explanations, which address "what if" scen…
counterfactualreinforcement-learningReinforcement LearningReinforcement Learning (RL)ACTER: Diverse and Actionable Counterfactual Sequences for Explaining and Diagnosing RL Policies
Understanding how failure occurs and how it can be prevented in reinforcement learning (RL) is necessary to enable debugging, maintain user trust, and develop personalized policies. Counterfactual reasoning has often bee…
counterfactualCounterfactual ReasoningDiversityreinforcement-learning+1Causal Discovery and Counterfactual Reasoning to Optimize Persuasive Dialogue Policies
Tailoring persuasive conversations to users leads to more effective persuasion. However, existing dialogue systems often struggle to adapt to dynamically evolving user states. This paper presents a novel method that leve…
Causal DiscoverycounterfactualCounterfactual Reasoning