paper-with-me

홈 › Papers

Rebounding Bandits for Modeling Satiation Effects

2020-11-13 · NeurIPS 2021 12 · Liu Leqi, Fatma Kilinc-Karzan, Zachary C. Lipton, Alan L. Montgomery

Psychological research shows that enjoyment of many goods is subject to satiation, with short-term satisfaction declining after repeated exposures to the same item. Nevertheless, proposed algorithms for powering recommender systems seldom model these dynamics, instead proceeding as though user preferences were fixed in time. In this work, we introduce rebounding bandits, a multi-armed bandit setup, where satiation dynamics are modeled as time-invariant linear dynamical systems. Expected rewards for each arm decline monotonically with consecutive exposures to it and rebound towards the initial reward whenever that arm is not pulled. Unlike classical bandit settings, methods for tackling rebounding bandits must plan ahead and model-based methods rely on estimating the parameters of the satiation dynamics. We characterize the planning problem, showing that the greedy policy is optimal when the arms exhibit identical deterministic dynamics. To address stochastic satiation dynamics with unknown parameters, we propose Explore-Estimate-Plan (EEP), an algorithm that pulls arms methodically, estimates the system dynamics, and then plans accordingly.

📄 PDF Abstract BibTeX arXiv:2011.06741

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation Systems

Similar Papers 제목 키워드 기반

A Last Switch Dependent Analysis of Satiation and Seasonality in Bandits

2021-10-22 · Pierre Laforgue, Giulia Clerici, Nicolò Cesa-Bianchi, Ran Gilad-Bachrach

Motivated by the fact that humans like some level of unpredictability or novelty, and might therefore get quickly bored when interacting with a stationary policy, we introduce a novel non-stationary bandit problem, where…

Linear Bandits with Memory: from Rotting to Rising

2023-02-16 · Giulia Clerici, Pierre Laforgue, Nicolò Cesa-Bianchi

Nonstationary phenomena, such as satiation effects in recommendations, have mostly been modeled using bandits with finitely many arms. However, the richer action space provided by linear bandits is often preferred in pra…

Decision MakingModel Selection

Last Switch Dependent Bandits with Monotone Payoff Functions

2023-06-01 · Ayoub Foussoul, Vineet Goyal, Orestis Papadigenopoulos, Assaf Zeevi

In a recent work, Laforgue et al. introduce the model of last switch dependent (LSD) bandits, in an attempt to capture nonstationary phenomena induced by the interaction between the player and the environment. Examples i…

Spillover Effects of US Monetary Policy on Emerging Markets Amidst Uncertainty

2024-02-11 · Povilas Lastauskas, Anh Dinh Minh Nguyen

This paper examines the impact of US monetary policy tightening on emerging markets, distinguishing between direct and indirect spillover effects using the global vector autoregression with stochastic volatility covering…

Contextual Multi-Armed Bandits for Causal Marketing

2018-10-02 · Neela Sawant, Chitti Babu Namballa, Narayanan Sadagopan, Houssam Nassif

This work explores the idea of a causal contextual multi-armed bandit approach to automated marketing, where we estimate and optimize the causal (incremental) effects. Focusing on causal effect leads to better return on …

Causal InferencecounterfactualMarketingMulti-Armed Bandits+1