paper-with-me

Papers

Density Matching Reward Learning

2016-08-12 · Sungjoon Choi, Kyungjae Lee, Andy Park, Songhwai Oh

In this paper, we focus on the problem of inferring the underlying reward function of an expert given demonstrations, which is often referred to as inverse reinforcement learning (IRL). In particular, we propose a model-free density-based IRL algorithm, named density matching reward learning (DMRL), which does not require model dynamics. The performance of DMRL is analyzed theoretically and the sample complexity is derived. Furthermore, the proposed DMRL is extended to handle nonlinear IRL problems by assuming that the reward function is in the reproducing kernel Hilbert space (RKHS) and kernel DMRL (KDMRL) is proposed. The parameters for KDMRL can be computed analytically, which greatly reduces the computation time. The performance of KDMRL is extensively evaluated in two sets of experiments: grid world and track driving experiments. In grid world experiments, the proposed KDMRL method is compared with both model-based and model-free IRL methods and shows superior performance on a nonlinear reward setting and competitive performance on a linear reward setting in terms of expected value differences. Then we move on to more realistic experiments of learning different driving styles for autonomous navigation in complex and dynamic tracks using KDMRL and receding horizon control.

📄 PDF Abstract BibTeX arXiv:1608.03694

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous NavigationReinforcement Learning

Similar Papers 제목 키워드 기반

f-IRL: Inverse Reinforcement Learning via State Marginal Matching

2020-11-09 · Tianwei Ni, Harshit Sikchi, YuFei Wang, Tejus Gupta 외

Imitation learning is well-suited for robotic tasks where it is difficult to directly program the behavior or specify a cost for optimal control. In this work, we propose a method for learning the reward function (and th…

Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

IL-flOw: Imitation Learning from Observation using Normalizing Flows

2022-05-19 · Wei-Di Chang, Juan Camilo Gamboa Higuera, Scott Fujimoto, David Meger 외

We present an algorithm for Inverse Reinforcement Learning (IRL) from expert state observations only. Our approach decouples reward modelling from policy learning, unlike state-of-the-art adversarial methods which requir…

continuous-controlContinuous ControlImitation Learningreinforcement-learning+2

Reinforcement Learning for Flow-Matching Policies with Density Transport

2026-06-07 · Boshu Lei, Kostas Daniilidis, Antonio Loquercio arxiv

We present an online reinforcement learning (RL) algorithm for fine-tuning flow-matching policies in continuous-control problems. Our key insight is to view RL-based policy improvement as a transport of action densities …

Reinforcement LearningRobot Manipulation

Tight Lower Bounds for the Multi-Secretary Problem via Bellman Certificates

2026-07-02 · Jiawei Zhang arxiv

This paper studies additive regret in the multi-secretary problem, defined as the gap between the expected offline prophet reward and the reward of the best online policy. Prior work established \(O(\log T)\) regret for …

A unified perspective on fine-tuning and sampling with diffusion and flow models

2026-04-30 · Carles Domingo-Enrich, Yuanqi Du, Michael S. Albergo arxiv

We study the problem of training diffusion and flow generative models to sample from target distributions defined by an exponential tilting of a base density; a formulation that subsumes both sampling from unnormalized d…