paper-with-me

홈 › Papers

Off-policy Confidence Sequences

2021-02-18 · Nikos Karampatziakis, Paul Mineiro, Aaditya Ramdas

We develop confidence bounds that hold uniformly over time for off-policy evaluation in the contextual bandit setting. These confidence sequences are based on recent ideas from martingale analysis and are non-asymptotic, non-parametric, and valid at arbitrary stopping times. We provide algorithms for computing these confidence sequences that strike a good balance between computational and statistical efficiency. We empirically demonstrate the tightness of our approach in terms of failure probability and width and apply it to the "gated deployment" problem of safely upgrading a production contextual bandit system.

📄 PDF Abstract BibTeX arXiv:2102.09540

Code (0)

등록된 구현이 없습니다.

Tasks

Off-policy evaluationvalid

Similar Papers 제목 키워드 기반

Improved Off-policy Reinforcement Learning in Biological Sequence Design

2024-10-06 · Hyeonah Kim, Minsu Kim, Taeyoung Yun, Sanghyeok Choi 외

Designing biological sequences with desired properties is challenging due to vast search spaces and limited evaluation budgets. Although reinforcement learning methods use proxy models for rapid reward evaluation, insuff…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies

2025-10-07 · Chunsan Hong, Seonho An, Min-Soo Kim, Jong Chul Ye arxiv

Masked diffusion models (MDMs) have recently emerged as a novel framework for language modeling. MDMs generate sentences by iteratively denoising masked sequences, filling in [MASK] tokens step by step. Although MDMs sup…

Semiparametric Efficient Inference in Adaptive Experiments

2023-11-30 · Thomas Cook, Alan Mishler, Aaditya Ramdas

We consider the problem of efficient inference of the Average Treatment Effect in a sequential experiment where the policy governing the assignment of subjects to treatment or control can change over time. We first provi…

valid

Catoni-style Confidence Sequences under Infinite Variance

2022-08-05 · Sujay Bhatt, Guanhua Fang, Ping Li, Gennady Samorodnitsky

In this paper, we provide an extension of confidence sequences for settings where the variance of the data-generating distribution does not exist or is infinite. Confidence sequences furnish confidence intervals that are…

valid

A new and flexible class of sharp asymptotic time-uniform confidence sequences

2025-02-14 · Felix Gnettner, Claudia Kirch

Confidence sequences are anytime-valid analogues of classical confidence intervals that do not suffer from multiplicity issues under optional continuation of the data collection. As in classical statistics, asymptotic co…

valid