paper-with-me

Papers

Case-based off-policy policy evaluation using prototype learning

2021-11-22 · Anton Matsson, Fredrik D. Johansson

Importance sampling (IS) is often used to perform off-policy policy evaluation but is prone to several issues, especially when the behavior policy is unknown and must be estimated from data. Significant differences between the target and behavior policies can result in uncertain value estimates due to, for example, high variance and non-evaluated actions. If the behavior policy is estimated using black-box models, it can be hard to diagnose potential problems and to determine for which inputs the policies differ in their suggested actions and resulting values. To address this, we propose estimating the behavior policy for IS using prototype learning. We apply this approach in the evaluation of policies for sepsis treatment, demonstrating how the prototypes give a condensed summary of differences between the target and behavior policies while retaining an accuracy comparable to baseline estimators. We also describe estimated values in terms of the prototypes to better understand which parts of the target policies have the most impact on the estimates. Using a simulator, we study the bias resulting from restricting models to use prototypes.

📄 PDF Abstract BibTeX arXiv:2111.11113

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Entity Graph Extraction from Legal Acts -- a Prototype for a Use Case in Policy Design Analysis

2022-09-02 · Anna Wróblewska, Bartosz Pieliński, Karolina Seweryn, Karol Saputa 외

This paper presents research on a prototype developed to serve the quantitative study of public policy design. This sub-discipline of political science focuses on identifying actors, relations between them, and tools at …

Online MDP with Transition Prototypes: A Robust Adaptive Approach

2024-12-18 · Shuo Sun, Meng Qi, Zuo-Jun Max Shen

In this work, we consider an online robust Markov Decision Process (MDP) where we have the information of finitely many prototypes of the underlying transition kernel. We consider an adaptively updated ambiguity set of t…

Decision MakingDecision Making Under Uncertainty

Balanced Off-Policy Evaluation for Personalized Pricing

2023-02-24 · Adam N. Elmachtoub, Vishal Gupta, Yunfan Zhao

We consider a personalized pricing problem in which we have data consisting of feature information, historical pricing decisions, and binary realized demand. The goal is to perform off-policy evaluation for a new persona…

Off-policy evaluation

Cross-Validated Off-Policy Evaluation

2024-05-24 · Matej Cief, Branislav Kveton, Michal Kompan

We study estimator selection and hyper-parameter tuning in off-policy evaluation. Although cross-validation is the most popular method for model selection in supervised learning, off-policy evaluation relies mostly on th…

Model SelectionOff-policy evaluation

A Semantic Framework for Enabling Radio Spectrum Policy Management and Evaluation

2020-11-08 · H. Santos, A. Mulvehill, J. S. Erickson, J. P. McCusker 외

Because radio spectrum is a finite resource, its usage and sharing is regulated by government agencies. These agencies define policies to manage spectrum allocation and assignment across multiple organizations, systems, …

Management