Orthogonal Policy Learning Under Ambiguity
This paper studies the problem of estimating individualized treatment rules when treatment effects are partially identified, as it is often the case with observational data. By drawing connections between the treatment assignment problem and classical decision theory, we characterize several notions of optimal treatment policies in the presence of partial identification. Our unified framework allows to incorporate user-defined constraints on the set of allowable policies, such as restrictions for transparency or interpretability, while also ensuring computational feasibility. We show how partial identification leads to a new policy learning problem where the objective function is directionally -- but not fully -- differentiable with respect to the nuisance first-stage. We then propose an estimation procedure that ensures Neyman-orthogonality with respect to the nuisance components and we provide statistical guarantees that depend on the amount of concentration around the points of non-differentiability in the data-generating-process. The proposed methods are illustrated using data from the Job Partnership Training Act study.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Singular Control in a Cash Management Model with Ambiguity
We consider a singular control model of cash reserve management, driven by a diffusion under ambiguity. The manager is assumed to have maxmin preferences over a set of priors characterized by $\kappa$-ignorance. A verifi…
ManagementFormalizing the Confluence of Orthogonal Rewriting Systems
Orthogonality is a discipline of programming that in a syntactic manner guarantees determinism of functional specifications. Essentially, orthogonality avoids, on the one side, the inherent ambiguity of non determinism, …
Identity-Expression Ambiguity in 3D Morphable Face Models
3D Morphable Models are a class of generative models commonly used to model faces. They are typically applied to ill-posed problems such as 3D reconstruction from 2D data. Several ambiguities in this problem's image form…
3D Reconstruction3D Shape GenerationInverse RenderingERPPO: Entropy Regularization-based Proximal Policy Optimization
Multi-Agent Proximal Policy Optimization (MAPPO) is a variant of the Proximal Policy Optimization (PPO) algorithm, specifically tailored for multi-agent reinforcement learning (MARL). MAPPO optimizes cooperative multi-ag…
Multi-agent Reinforcement LearningObject LocalizationObject DetectionWaveform Design for OFDM-based ISAC Systems Under Resource Occupancy Constraint
Integrated Sensing and Communication (ISAC) is one of the key pillars envisioned for 6G wireless systems. ISAC systems combine communication and sensing functionalities over a single waveform, with full resource sharing.…
Integrated sensing and communicationISACMatrix Completion