paper-with-me

Papers

Externally Valid Policy Evaluation Combining Trial and Observational Data

2023-10-23 · Sofia Ek, Dave Zachariah

Randomized trials are widely considered as the gold standard for evaluating the effects of decision policies. Trial data is, however, drawn from a population which may differ from the intended target population and this raises a problem of external validity (aka. generalizability). In this paper we seek to use trial data to draw valid inferences about the outcome of a policy on the target population. Additional covariate data from the target population is used to model the sampling of individuals in the trial study. We develop a method that yields certifiably valid trial-based policy evaluations under any specified range of model miscalibrations. The method is nonparametric and the validity is assured even with finite samples. The certified policy evaluations are illustrated using both simulated and real data.

📄 PDF Abstract BibTeX arXiv:2310.14763

Code (0)

등록된 구현이 없습니다.

Tasks

valid

Similar Papers 제목 키워드 기반

Trustworthy and Explainable Deep Reinforcement Learning for Safe and Energy-Efficient Process Control: A Use Case in Industrial Compressed Air Systems

2025-12-20 · Vincent Bezold, Patrick Wagner, Jakob Hofmann, Marco Huber 외 arxiv

This paper presents a trustworthy reinforcement learning approach for the control of industrial compressed air systems. We develop a framework that enables safe and energy-efficient operation under realistic boundary con…

Reinforcement Learning

Externally Valid Policy Choice

2022-05-11 · Christopher Adjaho, Timothy Christensen

We consider the problem of learning personalized treatment policies that are externally valid or generalizable: they perform well in other target populations besides the experimental (or training) population from which d…

valid

Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas

2025-09-26 · Luke Guerdan, Justin Whitehouse, Kimberly Truong, Kenneth Holstein 외 arxiv

As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-world deployment conditions. Threats to th…

Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation

2026-08-17 · Marc Pérez-Roig, David Fernández-Narro, Carlos Sáez arxiv

The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from histori…

Reinforcement Learning

Anytime-valid Optimal Policy Identification

2026-06-16 · Daniel Molitor arxiv

We develop an anytime-valid framework for optimal policy identification from logged contextual bandit data. In many applied settings, the analyst wants to select the optimal policy from a candidate policy class $Π$, but …