paper-with-me

Papers

Adaptive Hedging under Delayed Feedback

2019-02-27 · Alexander Korotin, Vladimir V'yugin, Evgeny Burnaev

The article is devoted to investigating the application of hedging strategies to online expert weight allocation under delayed feedback. As the main result, we develop the General Hedging algorithm $\mathcal{G}$ based on the exponential reweighing of experts' losses. We build the artificial probabilistic framework and use it to prove the adversarial loss bounds for the algorithm $\mathcal{G}$ in the delayed feedback setting. The designed algorithm $\mathcal{G}$ can be applied to both countable and continuous sets of experts. We also show how algorithm $\mathcal{G}$ extends classical Hedge (Multiplicative Weights) and adaptive Fixed Share algorithms to the delayed feedback and derive their regret bounds for the delayed setting by using our main result.

📄 PDF Abstract BibTeX arXiv:1902.10433

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Explicit Computations for Delayed Semistatic Hedging

2023-08-21 · Yan Dolinsky, Or Zuk

In this work we consider the exponential utility maximization problem in the framework of semistatic hedging.

Adaptive Experimentation with Delayed Binary Feedback

2022-02-02 · Zenan Wang, Carlos Carrion, Xiliang Lin, Fuhua Ji 외

Conducting experiments with objectives that take significant delays to materialize (e.g. conversions, add-to-cart events, etc.) is challenging. Although the classical "split sample testing" is still valid for the delayed…

Multi-Armed Banditsvalid

Delayed Bet-Hedging Resilience Strategies Under Environmental Fluctuations

2017-04-21

Many biological populations, such as bacterial colonies, have developed through evolution a protection mechanism, called bet-hedging, to increase their probability of survival under stressful environmental fluctutation. …

Deep Hedging with Market Impact

2024-02-20 · Andrei Neagu, Frédéric Godin, Clarence Simard, Leila Kosseim

Dynamic hedging is the practice of periodically transacting financial instruments to offset the risk caused by an investment or a liability. Dynamic hedging optimization can be framed as a sequential decision problem; th…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Multi-Armed Bandit Strategies for Non-Stationary Reward Distributions and Delayed Feedback Processes

2019-02-22 · Larkin Liu, Richard Downe, Joshua Reid

A survey is performed of various Multi-Armed Bandit (MAB) strategies in order to examine their performance in circumstances exhibiting non-stationary stochastic reward functions in conjunction with delayed feedback. We r…

Thompson Sampling