paper-with-me

홈 › Papers

Beyond Imitation: Recovering Dense Rewards from Demonstrations

2025-10-02 · Jiangnan Li, Thuy-Trang Vu, Ehsan Abbasnejad, Gholamreza Haffari arxiv

Conventionally, supervised fine-tuning (SFT) is treated as a simple imitation learning process that only trains a policy to imitate expert behavior on demonstration datasets. In this work, we challenge this view by establishing a fundamental equivalence between SFT and Inverse Reinforcement Learning. We prove that the SFT objective is a special case of Inverse Q-Learning, which implies that the SFT process does not just learn a policy, but also an implicit, dense, token-level reward model that explains the expert demonstrations. We then show how to recover this dense reward signal directly from the SFT model by formulating a baseline-relative reward function. The availability of such a dense reward model offers numerous benefits, providing granular credit assignment for each token generated. We demonstrate one key application by using these recovered rewards to further improve the policy with reinforcement learning. Our method, Dense-Path REINFORCE, consistently outperforms the original SFT models on instruction-following benchmarks. This work reframes SFT not merely as policy imitation but as a powerful reward learning mechanism, opening new possibilities for leveraging expert demonstrations.

📄 PDF Abstract BibTeX arXiv:2510.02493

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Visual Imitation Learning with Patch Rewards

2023-02-02 · Minghuan Liu, Tairan He, Weinan Zhang, Shuicheng Yan 외

Visual imitation learning enables reinforcement learning agents to learn to behave from expert visual demonstrations such as videos or image sequences, without explicit, well-defined rewards. Previous research either ado…

Imitation Learning

Wasserstein Adversarial Imitation Learning

2019-06-19 · Huang Xiao, Michael Herman, Joerg Wagner, Sebastian Ziesche 외

Imitation Learning describes the problem of recovering an expert policy from demonstrations. While inverse reinforcement learning approaches are known to be very sample-efficient in terms of expert demonstrations, they u…

Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration

2025-06-25 · Heyang Zhao, Xingrui Yu, David M. Bossens, Ivor W. Tsang 외

Imitation learning is a central problem in reinforcement learning where the goal is to learn a policy that mimics the expert's behavior. In practice, it is often challenging to learn the expert policy from a limited numb…

Imitation LearningMuJoCo

WoMAP: World Models For Embodied Open-Vocabulary Object Localization

2025-06-02 · Tenny Yin, Zhiting Mei, Tao Sun, Lihan Zha 외

Language-instructed active object localization is a critical challenge for robots, requiring efficient exploration of partially observable environments. However, state-of-the-art approaches either struggle to generalize …

Active Object LocalizationEfficient ExplorationImitation LearningObject+1

Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations

2019-07-09 · Daniel S. Brown, Wonjoon Goo, Scott Niekum

The performance of imitation learning is typically upper-bounded by the performance of the demonstrator. While recent empirical results demonstrate that ranked demonstrations allow for better-than-demonstrator performanc…

Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)