paper-with-me

홈 › Papers

Non-Stationary Bandits with Intermediate Observations

2020-01-01 · ICML 2020 1 · Claire Vernade, András György, Timothy Mann

Online recommender systems often face long delays in receiving feedback, especially when optimizing for some long-term metrics. While mitigating the effects of delays in learning is well-understood in stationary environments, the problem becomes much more challenging when the environment changes. In fact, if the timescale of the change is comparable to the delay, it is impossible to learn about the environment, since the available observations are already obsolete. However, the arising issues can be addressed if intermediate signals are available without delay, such that given those signals, the long-term behavior of the system is stationary. To model this situation, we introduce the problem of stochastic, non-stationary, delayed bandits with intermediate observations. We develop a computationally efficient algorithm based on $\UCRL$, and prove sublinear regret guarantees for its performance. Experimental results demonstrate that our method is able to learn in non-stationary delayed environments where existing methods fail.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation Systems

Similar Papers 제목 키워드 기반

Non-Stationary Delayed Bandits with Intermediate Observations

2020-06-03 · Claire Vernade, Andras Gyorgy, Timothy Mann

Online recommender systems often face long delays in receiving feedback, especially when optimizing for some long-term metrics. While mitigating the effects of delays in learning is well-understood in stationary environm…

Recommendation Systems

A Definition of Non-Stationary Bandits

2023-02-23 · Yueyang Liu, Xu Kuang, Benjamin Van Roy

Despite the subject of non-stationary bandit learning having attracted much recent attention, we have yet to identify a formal definition of non-stationarity that can consistently distinguish non-stationary bandits from …

Delayed Bandits: When Do Intermediate Observations Help?

2023-05-30 · Emmanuel Esposito, Saeed Masoudian, Hao Qiu, Dirk van der Hoeven 외

We study a $K$-armed bandit with delayed feedback and intermediate observations. We consider a model where intermediate observations have a form of a finite state, which is observed immediately after taking an action, wh…

Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent

2026-05-25 · Joongkyu Lee, Min-hwan Oh arxiv

We study nonstationary generalized linear bandits (GLBs), where the expected reward is modeled through a nonlinear link function with an unknown time-varying parameter. This framework encompasses a broad class of reward …

Computational Efficiency

Taming Non-stationary Bandits: A Bayesian Approach

2017-07-31 · Vishnu Raj, Sheetal Kalyani

We consider the multi armed bandit problem in non-stationary environments. Based on the Bayesian method, we propose a variant of Thompson Sampling which can be used in both rested and restless bandit scenarios. Applying …

Thompson Sampling