paper-with-me

Papers

Semi-Parametric Efficient Policy Learning with Continuous Actions

2019-05-24 · NeurIPS 2019 12 · Mert Demirer, Vasilis Syrgkanis, Greg Lewis, Victor Chernozhukov

We consider off-policy evaluation and optimization with continuous action spaces. We focus on observational data where the data collection policy is unknown and needs to be estimated. We take a semi-parametric approach where the value function takes a known parametric form in the treatment, but we are agnostic on how it depends on the observed contexts. We propose a doubly robust off-policy estimate for this setting and show that off-policy optimization based on this estimate is robust to estimation errors of the policy function or the regression model. Our results also apply if the model does not satisfy our semi-parametric form, but rather we measure regret in terms of the best projection of the true value function to this functional space. Our work extends prior approaches of policy optimization from observational data that only considered discrete actions. We provide an experimental evaluation of our method in a synthetic data example motivated by optimal personalized pricing and costly resource allocation.

📄 PDF Abstract BibTeX arXiv:1905.10116

Code (1)

vsyrgkanis/policy_learning_continuous_actions 공식 구현

Tasks

Off-policy evaluation

Similar Papers 제목 키워드 기반

Difference-Aware Retrieval Policies for Imitation Learning

2026-06-08 · Quinn Pfeifer, Ethan Pronovost, Paarth Shah, Khimya Khetarpal 외 arxiv

Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment. We show that reusing the training data during inference vi…

Continuous Control

Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models

2025-12-30 · Lars van der Laan, Aurelien Bibaut, Nathan Kallus arxiv

In many sequential decision-making problems, researchers observe actions but not the rewards that drive behavior, yet still wish to evaluate and compare counterfactual policies. Inverse reinforcement learning (IRL) and d…

Reinforcement Learning

Solving the non-preemptive two queue polling model with generally distributed service and switch-over durations and Poisson arrivals as a Semi-Markov Decision Process

2021-12-13 · Dylan Solms

The polling system with switch-over durations is a useful model with several practical applications. It is classified as a Discrete Event Dynamic System (DEDS) for which no one agreed upon modelling approach exists. Furt…

Automatic Double Reinforcement Learning in Semiparametric Markov Decision Processes with Applications to Long-Term Causal Inference

2025-01-12 · Lars van der Laan, David Hubbard, Allen Tran, Nathan Kallus 외

Estimating long-term causal effects from short-term data is essential for decision-making in healthcare, economics, and industry, where long-term follow-up is often infeasible. Markov Decision Processes (MDPs) offer a pr…

Causal InferenceDimensionality ReductionDomain AdaptationModel Selection

On Well-posedness and Minimax Optimal Rates of Nonparametric Q-function Estimation in Off-policy Evaluation

2022-01-17 · Xiaohong Chen, Zhengling Qi

We study the off-policy evaluation (OPE) problem in an infinite-horizon Markov decision process with continuous states and actions. We recast the $Q$-function estimation into a special form of the nonparametric instrumen…

Off-policy evaluation