Adaptive Hedging under Delayed Feedback
The article is devoted to investigating the application of hedging strategies to online expert weight allocation under delayed feedback. As the main result, we develop the General Hedging algorithm $\mathcal{G}$ based on the exponential reweighing of experts' losses. We build the artificial probabilistic framework and use it to prove the adversarial loss bounds for the algorithm $\mathcal{G}$ in the delayed feedback setting. The designed algorithm $\mathcal{G}$ can be applied to both countable and continuous sets of experts. We also show how algorithm $\mathcal{G}$ extends classical Hedge (Multiplicative Weights) and adaptive Fixed Share algorithms to the delayed feedback and derive their regret bounds for the delayed setting by using our main result.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Explicit Computations for Delayed Semistatic Hedging
In this work we consider the exponential utility maximization problem in the framework of semistatic hedging.
Adaptive Experimentation with Delayed Binary Feedback
Conducting experiments with objectives that take significant delays to materialize (e.g. conversions, add-to-cart events, etc.) is challenging. Although the classical "split sample testing" is still valid for the delayed…
Multi-Armed BanditsvalidDelayed Bet-Hedging Resilience Strategies Under Environmental Fluctuations
Many biological populations, such as bacterial colonies, have developed through evolution a protection mechanism, called bet-hedging, to increase their probability of survival under stressful environmental fluctutation. …
Deep Hedging with Market Impact
Dynamic hedging is the practice of periodically transacting financial instruments to offset the risk caused by an investment or a liability. Dynamic hedging optimization can be framed as a sequential decision problem; th…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Multi-Armed Bandit Strategies for Non-Stationary Reward Distributions and Delayed Feedback Processes
A survey is performed of various Multi-Armed Bandit (MAB) strategies in order to examine their performance in circumstances exhibiting non-stationary stochastic reward functions in conjunction with delayed feedback. We r…
Thompson Sampling