Optimal Decision Rules when Payoffs are Partially Identified
We derive asymptotically optimal statistical decision rules for discrete choice problems when payoffs depend on a partially-identified parameter $\theta$ and the decision maker can use a point-identified parameter $\mu$ to deduce restrictions on $\theta$. Examples include treatment choice under partial identification and pricing with rich unobserved heterogeneity. Our notion of optimality combines a minimax approach to handle the ambiguity from partial identification of $\theta$ given $\mu$ with an average risk minimization approach for $\mu$. We show how to implement optimal decision rules using the bootstrap and (quasi-)Bayesian methods in both parametric and semiparametric settings. We provide detailed applications to treatment choice and optimal pricing. Our asymptotic approach is well suited for realistic empirical settings in which the derivation of finite-sample optimal rules is intractable.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Partially Observable Contextual Bandits with Linear Payoffs
The standard contextual bandit framework assumes fully observable and actionable contexts. In this work, we consider a new bandit setting with partially observable, correlated contexts and linear payoffs, motivated by th…
Decision MakingMulti-Armed Banditsparameter estimationThompson SamplingSelling Information
I consider the monopolistic pricing of informational good. A buyer's willingness to pay for information is from inferring the unknown payoffs of actions in decision making. A monopolistic seller and the buyer each observ…
Decision MakingThreshold UCT: Cost-Constrained Monte Carlo Tree Search with Pareto Curves
Constrained Markov decision processes (CMDPs), in which the agent optimizes expected payoffs while keeping the expected cost below a given threshold, are the leading framework for safe sequential decision making under st…
Decision MakingSequential Decision MakingTreatment Choice, Mean Square Regret and Partial Identification
We consider a decision maker who faces a binary treatment choice when their welfare is only partially identified from data. We contribute to the literature by anchoring our finite-sample analysis on mean square regret, a…
Approximate optimality and the risk/reward tradeoff in a class of bandit problems
This paper studies a sequential decision problem where payoff distributions are known and where the riskiness of payoffs matters. Equivalently, it studies sequential choice from a repeated set of independent lotteries. T…