paper-with-me

홈 › Papers

Learning to Use Learners' Advice

2017-02-16 · Adish Singla, Hamed Hassani, Andreas Krause

In this paper, we study a variant of the framework of online learning using expert advice with limited/bandit feedback. We consider each expert as a learning entity, seeking to more accurately reflecting certain real-world applications. In our setting, the feedback at any time $t$ is limited in a sense that it is only available to the expert $i^t$ that has been selected by the central algorithm (forecaster), \emph{i.e.}, only the expert $i^t$ receives feedback from the environment and gets to learn at time $t$. We consider a generic black-box approach whereby the forecaster does not control or know the learning dynamics of the experts apart from knowing the following no-regret learning property: the average regret of any expert $j$ vanishes at a rate of at least $O(t_j^{\regretRate-1})$ with $t_j$ learning steps where $\regretRate \in [0, 1]$ is a parameter. In the spirit of competing against the best action in hindsight in multi-armed bandits problem, our goal here is to be competitive w.r.t. the cumulative losses the algorithm could receive by following the policy of always selecting one expert. We prove the following hardness result: without any coordination between the forecaster and the experts, it is impossible to design a forecaster achieving no-regret guarantees. In order to circumvent this hardness result, we consider a practical assumption allowing the forecaster to "guide" the learning process of the experts by filtering/blocking some of the feedbacks observed by them from the environment, \emph{i.e.}, not allowing the selected expert $i^t$ to learn at time $t$ for some time steps. Then, we design a novel no-regret learning algorithm \algo for this problem setting by carefully guiding the feedbacks observed by experts. We prove that \algo achieves the worst-case expected cumulative regret of $O(\Time^\frac{1}{2 - \regretRate})$ after $\Time$ time steps.

📄 PDF Abstract BibTeX arXiv:1702.04825

Code (0)

등록된 구현이 없습니다.

Tasks

BlockingMulti-Armed Bandits

Similar Papers 제목 키워드 기반

Optimal Prediction Using Expert Advice and Randomized Littlestone Dimension

2023-02-27 · Yuval Filmus, Steve Hanneke, Idan Mehalel, Shay Moran

A classical result in online learning characterizes the optimal mistake bound achievable by deterministic learners using the Littlestone dimension (Littlestone '88). We prove an analogous result for randomized learners: …

2kOpen-Ended Question Answering

Multi-Advisor Reinforcement Learning

2017-04-03 · ICLR 2018 1 · Romain Laroche, Mehdi Fatemi, Joshua Romoff, Harm van Seijen

We consider tackling a single-agent RL problem by distributing it to $n$ learners. These learners, called advisors, endeavour to solve the problem from a different focus. Their advice, taking the form of action values, i…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Application of Deep Self-Attention in Knowledge Tracing

2021-05-17 · Junhao Zeng, Qingchun Zhang, Ning Xie, Bochun Yang

The development of intelligent tutoring system has greatly influenced the way students learn and practice, which increases their learning efficiency. The intelligent tutoring system must model learners' mastery of the kn…

Knowledge Tracing

A Broad-persistent Advising Approach for Deep Interactive Reinforcement Learning in Robotic Environments

2021-10-15 · Hung Son Nguyen, Francisco Cruz, Richard Dazeley

Deep Reinforcement Learning (DeepRL) methods have been widely used in robotics to learn about the environment and acquire behaviors autonomously. Deep Interactive Reinforcement Learning (DeepIRL) includes interactive fee…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

A Prescriptive Learning Analytics Framework: Beyond Predictive Modelling and onto Explainable AI with Prescriptive Analytics and ChatGPT

2022-08-31 · Teo Susnjak

A significant body of recent research in the field of Learning Analytics has focused on leveraging machine learning approaches for predicting at-risk students in order to initiate timely interventions and thereby elevate…