paper-with-me

홈 › Papers

Learning Safely Without Knowing the World:COMPASS-Hedge

2026-03-22 · Ting Hu, Luanda Cai, Emmanouil-Vasileios Vlatakis-Gkaragkounis arxiv

Online learning algorithms often face a fundamental trilemma: balancing regret guarantees between adversarial and stochastic settings and providing baseline safety against a fixed comparator. While existing methods excel in one or two of these regimes, they typically fail to unify all three without sacrificing optimal rates or requiring oracle access to problem-dependent parameters. In this work, we bridge this gap by introducing COMPASS-Hedge. To the best of our knowledge, our algorithm is the first full-information anytime method to simultaneously achieve, up to logarithmic factors: i) minimax-optimal regret in adversarial environments; ii) instance-optimal, gap-dependent regret in stochastic environments; and iii) $\tilde{\mathcal{O}}(1)$ regret relative to a designated baseline policy. Crucially, COMPASS-Hedge is parameter-free and requires no prior knowledge of the environment's nature or the magnitude of the stochastic suboptimality gaps. Our approach hinges on a novel integration of adaptive pseudo-regret scaling and phase-based aggression, coupled with a comparator-aware mixing strategy. To the best of our knowledge, this provides the first "best-of-three-world" guarantee in the full-information setting, establishing that baseline safety does not have to come at the cost of worst-case robustness or stochastic efficiency.

📄 PDF Abstract BibTeX arXiv:2603.22348

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Statistical Arbitrage Risk Premium by Machine Learning

2021-03-18 · Raymond C. W. Leung, Yu-Man Tam

How to hedge factor risks without knowing the identities of the factors? We first prove a general theoretical result: even if the exact set of factors cannot be identified, any risky asset can use some portfolio of simil…

BIG-bench Machine LearningPosition

Using Hedge Detection to Improve Committed Belief Tagging

2018-06-01 · WS 2018 6 · Morgan Ulinski, Seth Benjamin, Julia Hirschberg

We describe a novel method for identifying hedge terms using a set of manually constructed rules. We present experiments adding hedge features to a committed belief system to improve classification. We compare performanc…

General ClassificationSentence Classification

Wildfire Smoke and Air Quality: How Machine Learning Can Guide Forest Management

2020-10-09 · Lorenzo Tomaselli, Coty Jen, Ann B. Lee

Prescribed burns are currently the most effective method of reducing the risk of widespread wildfires, but a largely missing component in forest management is knowing which fuels one can safely burn to minimize exposure …

BIG-bench Machine LearningClusteringManagement

Knowing Whether

2013-11-30 · Jie Fan, Yanjing Wang, Hans van Ditmarsch

Knowing whether a proposition is true means knowing that it is true or knowing that it is false. In this paper, we study logics with a modal operator Kw for knowing whether but without a modal operator K for knowing that…

QLBS: Q-Learner in the Black-Scholes(-Merton) Worlds

2017-12-13 · Igor Halperin

This paper presents a discrete-time option pricing model that is rooted in Reinforcement Learning (RL), and more specifically in the famous Q-Learning method of RL. We construct a risk-adjusted Markov Decision Process fo…

BenchmarkingModel-based Reinforcement LearningQ-Learningreinforcement-learning+2