paper-with-me

홈 › Papers

Statistical Inference for Misspecified Contextual Bandits

2025-09-08 · Yongyi Guo, Ziping Xu arxiv

Contextual bandit algorithms have transformed modern experimentation by enabling real-time adaptation for personalized treatment and efficient use of data. Yet these advantages create challenges for statistical inference due to adaptivity. A fundamental property that supports valid inference is policy convergence, meaning that action-selection probabilities converge in probability given the context. Convergence ensures replicability of adaptive experiments and stability of online algorithms. In this paper, we highlight a previously overlooked issue: widely used algorithms such as LinUCB may fail to converge when the reward model is misspecified, and such non-convergence creates fundamental obstacles for statistical inference. This issue is practically important, as misspecified models -- such as linear approximations of complex dynamic system -- are often employed in real-world adaptive experiments to balance bias and variance. Motivated by this insight, we propose and analyze a broad class of algorithms that are guaranteed to converge even under model misspecification. Building on this guarantee, we develop a general inference framework based on an inverse-probability-weighted Z-estimator (IPW-Z) and establish its asymptotic normality with a consistent variance estimator. Simulation studies confirm that the proposed method provides robust and data-efficient confidence intervals, and can outperform existing approaches that exist only in the special case of offline policy evaluation. Taken together, our results underscore the importance of designing adaptive algorithms with built-in convergence guarantees to enable stable experimentation and valid statistical inference in practice.

📄 PDF Abstract BibTeX arXiv:2509.06287

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Statistical Inference for Misspecified Contextual Bandits

2026-06-21 · Yongyi Guo, Ziping Xu arxiv

Contextual bandit algorithms have transformed modern experimentation by enabling real-time adaptation for personalized treatment. Yet these advantages create challenges for statistical inference due to adaptivity. We stu…

Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs

2023-10-13 · Debangshu Banerjee, Aditya Gopalan

Parametric, feature-based reward models are employed by a variety of algorithms in decision-making settings such as bandits and Markov decision processes (MDPs). The typical assumption under which the algorithms are anal…

Decision MakingMulti-Armed BanditsQ-Learning

Bayesian decision-making under misspecified priors with applications to meta-learning

2021-07-03 · NeurIPS 2021 12 · Max Simchowitz, Christopher Tosh, Akshay Krishnamurthy, Daniel Hsu 외

Thompson sampling and other Bayesian sequential decision-making algorithms are among the most popular approaches to tackle explore/exploit trade-offs in (contextual) bandits. The choice of prior in these algorithms offer…

Decision MakingMeta-LearningMulti-Armed BanditsSequential Decision Making+1

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification

2026-06-04 · Haoyang Hong, Zichen Wang, Quanquan Gu, Huazheng Wang arxiv

We study KL-regularized contextual bandits and episodic reinforcement learning (RL) under general function approximation with model misspecification. Existing guarantees rely on realizability and therefore do not extend …

Reinforcement Learning

Adapting to Misspecification in Contextual Bandits

2021-07-12 · NeurIPS 2020 12 · Dylan J. Foster, Claudio Gentile, Mehryar Mohri, Julian Zimmert

A major research direction in contextual bandits is to develop algorithms that are computationally efficient, yet support flexible, general-purpose function approximation. Algorithms based on modeling rewards have shown …

Multi-Armed Banditsregression