paper-with-me

홈 › Papers

Apprenticeship learning with prior beliefs using inverse optimization

2025-05-27 · Mauricio Junca, Esteban Leiva

The relationship between inverse reinforcement learning (IRL) and inverse optimization (IO) for Markov decision processes (MDPs) has been relatively underexplored in the literature, despite addressing the same problem. In this work, we revisit the relationship between the IO framework for MDPs, IRL, and apprenticeship learning (AL). We incorporate prior beliefs on the structure of the cost function into the IRL and AL problems, and demonstrate that the convex-analytic view of the AL formalism (Kamoutsi et al., 2021) emerges as a relaxation of our framework. Notably, the AL formalism is a special case in our framework when the regularization term is absent. Focusing on the suboptimal expert setting, we formulate the AL problem as a regularized min-max problem. The regularizer plays a key role in addressing the ill-posedness of IRL by guiding the search for plausible cost functions. To solve the resulting regularized-convex-concave-min-max problem, we use stochastic mirror descent (SMD) and establish convergence bounds for the proposed method. Numerical experiments highlight the critical role of regularization in learning cost vectors and apprentice policies.

📄 PDF Abstract BibTeX arXiv:2505.21639

Code (1)

EstebanLeiva/apprenticeshiplearning 공식 구현

Similar Papers 제목 키워드 기반

Stochastic convex optimization for provably efficient apprenticeship learning

2021-12-31 · Angeliki Kamoutsi, Goran Banjac, John Lygeros

We consider large-scale Markov decision processes (MDPs) with an unknown cost function and employ stochastic convex optimization tools to address the problem of imitation learning, which consists of learning a policy fro…

Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Interpretable Apprenticeship Learning with Temporal Logic Specifications

2017-10-28 · Daniel Kasenberg, Matthias Scheutz

Recent work has addressed using formulas in linear temporal logic (LTL) as specifications for agents planning in Markov Decision Processes (MDPs). We consider the inverse problem: inferring an LTL specification from demo…

Multiobjective Optimization

Safety-Aware Multi-Agent Apprenticeship Learning

2022-01-20 · Junchen Zhao

Our objective of this project is to make the extension based on the technique mentioned in the paper "Safety-Aware Apprenticeship Learning" to improve the utility and the efficiency of the existing Reinforcement Learning…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Online Apprenticeship Learning

2021-02-13 · Lior Shani, Tom Zahavy, Shie Mannor

In Apprenticeship Learning (AL), we are given a Markov Decision Process (MDP) without access to the cost function. Instead, we observe trajectories sampled by an expert that acts according to some policy. The goal is to …

Bootstrapping Apprenticeship Learning

2010-12-01 · NeurIPS 2010 12 · Abdeslam Boularias, Brahim Chaib-Draa

We consider the problem of apprenticeship learning where the examples, demonstrated by an expert, cover only a small part of a large state space. Inverse Reinforcement Learning (IRL) provides an efficient tool for genera…

Car RacingReinforcement Learning