paper-with-me

Papers

One Arrow, Two Kills: An Unified Framework for Achieving Optimal Regret Guarantees in Sleeping Bandits

2022-10-26 · Pierre Gaillard, Aadirupa Saha, Soham Dan

We address the problem of \emph{`Internal Regret'} in \emph{Sleeping Bandits} in the fully adversarial setup, as well as draw connections between different existing notions of sleeping regrets in the multiarmed bandits (MAB) literature and consequently analyze the implications: Our first contribution is to propose the new notion of \emph{Internal Regret} for sleeping MAB. We then proposed an algorithm that yields sublinear regret in that measure, even for a completely adversarial sequence of losses and availabilities. We further show that a low sleeping internal regret always implies a low external regret, and as well as a low policy regret for iid sequence of losses. The main contribution of this work precisely lies in unifying different notions of existing regret in sleeping bandits and understand the implication of one to another. Finally, we also extend our results to the setting of \emph{Dueling Bandits} (DB)--a preference feedback variant of MAB, and proposed a reduction to MAB idea to design a low regret algorithm for sleeping dueling bandits with stochastic preferences and adversarial availabilities. The efficacy of our algorithms is justified through empirical evaluations.

📄 PDF Abstract BibTeX arXiv:2210.14998

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Framework for Searching for General Artificial Intelligence

2016-11-02 · Marek Rosa, Jan Feyereisl, The GoodAI Collective

There is a significant lack of unified approaches to building generally intelligent machines. The majority of current artificial intelligence research operates within a very narrow field of focus, frequently without cons…

Progressively Learning Heterogeneous Skills in a Unified Latent Space

2026-08-24 · Yue-Yi Zhang, Ming Gong, Linpu He, Wei-Shi Zheng 외 arxiv

We propose HetSkills, a novel framework designed to progressively learn heterogeneous skills within a unified latent space for physics-based character control. The core idea is to treat this latent space as a shared exec…

On the creation of narrow AI: hierarchy and nonlocality of neural network skills

2025-05-21 · Eric J. Michaud, Asher Parker-Sartori, Max Tegmark

We study the problem of creating strong, yet narrow, AI systems. While recent AI progress has been driven by the training of large general-purpose foundation models, the creation of smaller models specialized for narrow …

HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos

2026-02-02 · Yinhuai Wang, Qihan Zhao, Yuen Fui Lau, Runyi Yu 외 arxiv

Enabling humanoid robots to perform agile and adaptive interactive tasks has long been a core challenge in robotics. Current approaches are bottlenecked by either the scarcity of realistic interaction data or the need fo…

Data Augmentation

Meta Context Engineering via Agentic Skill Evolution

2026-01-29 · Haoran Ye, Xuning He, Vincent Arak, Haonan Dong 외 arxiv

The operational efficacy of large language models relies heavily on their inference-time context. This has established Context Engineering (CE) as a formal discipline for optimizing these inputs. Current CE methods rely …