paper-with-me

홈 › Papers

Policy Optimization as Online Learning with Mediator Feedback

2020-12-15 · Alberto Maria Metelli, Matteo Papini, Pierluca D'Oro, Marcello Restelli

Policy Optimization (PO) is a widely used approach to address continuous control tasks. In this paper, we introduce the notion of mediator feedback that frames PO as an online learning problem over the policy space. The additional available information, compared to the standard bandit feedback, allows reusing samples generated by one policy to estimate the performance of other policies. Based on this observation, we propose an algorithm, RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST), for regret minimization in PO, that employs a randomized exploration strategy, differently from the existing optimistic approaches. When the policy space is finite, we show that under certain circumstances, it is possible to achieve constant regret, while always enjoying logarithmic regret. We also derive problem-dependent regret lower bounds. Then, we extend RANDOMIST to compact policy spaces. Finally, we provide numerical simulations on finite and compact policy spaces, in comparison with PO and bandit baselines.

📄 PDF Abstract BibTeX arXiv:2012.08225

Code (0)

등록된 구현이 없습니다.

Tasks

continuous-controlContinuous Control

Similar Papers 제목 키워드 기반

Pure Exploration under Mediators' Feedback

2023-08-29 · Riccardo Poiani, Alberto Maria Metelli, Marcello Restelli

Stochastic multi-armed bandits are a sequential-decision-making framework, where, at each interaction step, the learner selects an arm and observes a stochastic reward. Within the context of best-arm identification (BAI)…

Decision MakingMulti-Armed BanditsSequential Decision Making

Information Capacity Regret Bounds for Bandits with Mediator Feedback

2024-02-15 · Khaled Eldowa, Nicolò Cesa-Bianchi, Alberto Maria Metelli, Marcello Restelli

This work addresses the mediator feedback problem, a bandit game where the decision set consists of a number of policies, each associated with a probability distribution over a common space of outcomes. Upon choosing a p…

Online Decision Mediation

2023-10-28 · Daniel Jarrett, Alihan Hüyük, Mihaela van der Schaar

Consider learning a decision support assistant to serve as an intermediary between (oracle) expert behavior and (imperfect) human behavior: At each time, the algorithm observes an action chosen by a fallible agent, and d…

Decision MakingDescriptive

Practical Performative Policy Learning with Strategic Agents

2024-12-02 · Qianyi Chen, Ying Chen, Bo Li

This paper studies the performative policy learning problem, where agents adjust their features in response to a released policy to improve their potential outcomes, inducing an endogenous distribution shift. There has b…

Causal Inference

LLMediator: GPT-4 Assisted Online Dispute Resolution

2023-07-27 · Hannes Westermann, Jaromir Savelka, Karim Benyekhlef

In this article, we introduce LLMediator, an experimental platform designed to enhance online dispute resolution (ODR) by utilizing capabilities of state-of-the-art large language models (LLMs) such as GPT-4. In the cont…