Model-Free and Model-Based Policy Evaluation when Causality is Uncertain
When decision-makers can directly intervene, policy evaluation algorithms give valid causal estimates. In off-policy evaluation (OPE), there may exist unobserved variables that both impact the dynamics and are used by the unknown behavior policy. These "confounders" will introduce spurious correlations and naive estimates for a new policy will be biased. We develop worst-case bounds to assess sensitivity to these unobserved confounders in finite horizons when confounders are drawn iid each period. We demonstrate that a model-based approach with robust MDPs gives sharper lower bounds by exploiting domain knowledge about the dynamics. Finally, we show that when unobserved confounders are persistent over time, OPE is far more difficult and existing techniques produce extremely conservative bounds.
Code (1)
Tasks
modelOff-policy evaluationSensitivityvalidSimilar Papers 제목 키워드 기반
Confidence-aware 3D Gaze Estimation and Evaluation Metric
Deep learning appearance-based 3D gaze estimation is gaining popularity due to its minimal hardware requirements and being free of constraint. Unreliable and overconfident inferences, however, still limit the adoption of…
Gaze EstimationText as data: a machine learning-based approach to measuring uncertainty
The Economic Policy Uncertainty index had gained considerable traction with both academics and policy practitioners. Here, we analyse news feed data to construct a simple, general measure of uncertainty in the United Sta…
BIG-bench Machine LearningOn the Behavioral Consequences of Reverse Causality
Reverse causality is a common causal misperception that distorts the evaluation of private actions and public policies. This paper explores the implications of this error when a decision maker acts on it and therefore af…
ACE : Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
The varying significance of distinct primitive behaviors during the policy learning process has been overlooked by prior model-free RL algorithms. Leveraging this insight, we explore the causal relationship between diffe…
continuous-controlContinuous ControlEfficient ExplorationMultiscale Causal Analysis of Market Efficiency via News Uncertainty Networks and the Financial Chaos Index
This study evaluates the scale-dependent informational efficiency of stock markets using the Financial Chaos Index, a tensor-eigenvalue-based measure of realized volatility. Incorporating Granger causality and network-th…