paper-with-me

홈 › Papers

Conformal Policy Learning with Distribution-Free Safety Guarantees

2026-09-15 · Ying Jin, Naoki Egami arxiv

Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of ``do no harm.'' In this paper, we propose \textit{conformal policy learning} (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs.

📄 PDF Abstract BibTeX arXiv:2609.17296

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Robust Conformal CBF and CLF Controllers via Iterative Policy Updates

2026-06-13 · Omid Mirzaeedodangeh, Eliot Shekhtman, Nikolai Matni, Lars Lindemann arxiv

Conformal prediction (CP) has been used to obtain probabilistic bounds on the error between a learned dynamics model and the true but unknown system. Such CP bounds can then be embedded into robust control Lyapunov funct…

Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction

2025-11-13 · Omid Mirzaeedodangeh, Eliot Shekhtman, Nikolai Matni, Lars Lindemann arxiv

Safe planning of an autonomous agent in interactive environments -- such as the control of a self-driving vehicle among pedestrians -- poses a major challenge as the behavior of the environment is unknown and reactive to…

Learning to Navigate Under Imperfect Perception: Conformalised Segmentation for Safe Reinforcement Learning

2025-10-21 · Daniel Bethell, Simos Gerasimou, Radu Calinescu, Calum Imrie arxiv

Reliable navigation in safety-critical environments requires both accurate hazard perception and principled uncertainty handling to strengthen downstream safety handling. Despite the effectiveness of existing approaches,…

Reinforcement LearningSemantic Segmentation

Conformal Policy Learning for Sensorimotor Control Under Distribution Shifts

2023-11-02 · Huang Huang, Satvik Sharma, Antonio Loquercio, Anastasios Angelopoulos 외

This paper focuses on the problem of detecting and reacting to changes in the distribution of a sensorimotor controller's observables. The key idea is the design of switching policies that can take conformal quantiles as…

Autonomous DrivingConformal Prediction

Safe Control using Learned Safety Filters and Adaptive Conformal Inference

2026-04-20 · Sacha Huriot, Ihab Tabbara, Hussein Sibai arxiv

Safety filters have been shown to be effective tools to ensure the safety of control systems with unsafe nominal policies. To address scalability challenges in traditional synthesis methods, learning-based approaches hav…