paper-with-me

홈 › Papers

Reward Compatibility: A Framework for Inverse RL

2025-01-14 · Filippo Lazzati, Mirco Mutti, Alberto Metelli

We provide an original theoretical study of Inverse Reinforcement Learning (IRL) through the lens of reward compatibility, a novel framework to quantify the compatibility of a reward with the given expert's demonstrations. Intuitively, a reward is more compatible with the demonstrations the closer the performance of the expert's policy computed with that reward is to the optimal performance for that reward. This generalizes the notion of feasible reward set, the most common framework in the theoretical IRL literature, for which a reward is either compatible or not compatible. The grayscale introduced by the reward compatibility is the key to extend the realm of provably efficient IRL far beyond what is attainable with the feasible reward set: from tabular to large-scale MDPs. We analyze the IRL problem across various settings, including optimal and suboptimal expert's demonstrations and both online and offline data collection. For all of these dimensions, we provide a tractable algorithm and corresponding sample complexity analysis, as well as various insights on reward compatibility and how the framework can pave the way to yet more general problem settings.

📄 PDF Abstract BibTeX arXiv:2501.07996

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How does Inverse RL Scale to Large State Spaces? A Provably Efficient Approach

2024-06-06 · Filippo Lazzati, Mirco Mutti, Alberto Maria Metelli

In online Inverse Reinforcement Learning (IRL), the learner can collect samples about the dynamics of the environment to improve its estimate of the reward function. Since IRL suffers from identifiability issues, many th…

Reinforcement learning for inverse structural design and rapid laser cutting of kirigami prototypes

2026-04-16 · Milad Yazdani, Shahriar Shalileh, Dena Shahriari arxiv

Kirigami is an increasingly useful fabrication method to produce shape-programmable metamaterial structures. However, inverse design remains difficult because deployment is nonlinear, and feasible cut layouts must satisf…

Reinforcement Learning

Whole Brain Susceptibility Mapping Using Harmonic Incompatibility Removal

2018-05-31 · Chenglong Bao, Jae Kyu Choi, Bin Dong

Quantitative susceptibility mapping (QSM) aims to visualize the three dimensional susceptibility distribution by solving the field-to-source inverse problem using the phase data in magnetic resonance signal. However, the…

Bounded Risk-Sensitive Markov Games: Forward Policy Design and Inverse Reward Learning with Iterative Reasoning and Cumulative Prospect Theory

2020-09-03 · Ran Tian, Liting Sun, Masayoshi Tomizuka

Classical game-theoretic approaches for multi-agent systems in both the forward policy design problem and the inverse reward learning problem often make strong rationality assumptions: agents perfectly maximize expected …

Visual IRL for Human-Like Robotic Manipulation

2024-12-16 · Ehsan Asali, Prashant Doshi

We present a novel method for collaborative robots (cobots) to learn manipulation tasks and perform them in a human-like manner. Our method falls under the learn-from-observation (LfO) paradigm, where robots learn to per…