paper-with-me

홈 › Papers

Supervised Reward Inference

2025-02-25 · Will Schwarzer, Jordan Schneider, Philip S. Thomas, Scott Niekum

Existing approaches to reward inference from behavior typically assume that humans provide demonstrations according to specific models of behavior. However, humans often indicate their goals through a wide range of behaviors, from actions that are suboptimal due to poor planning or execution to behaviors which are intended to communicate goals rather than achieve them. We propose that supervised learning offers a unified framework to infer reward functions from any class of behavior, and show that such an approach is asymptotically Bayes-optimal under mild assumptions. Experiments on simulated robotic manipulation tasks show that our method can efficiently infer rewards from a wide variety of arbitrarily suboptimal demonstrations.

📄 PDF Abstract BibTeX arXiv:2502.18447

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-Supervised Online Reward Shaping in Sparse-Reward Environments

2021-03-08 · Farzan Memarian, Wonjoon Goo, Rudolf Lioutikov, Scott Niekum 외

We introduce Self-supervised Online Reward Shaping (SORS), which aims to improve the sample efficiency of any RL algorithm in sparse-reward environments by automatically densifying rewards. The proposed framework alterna…

On Designing Effective RL Reward at Training Time for LLM Reasoning

2024-10-19 · Jiaxuan Gao, Shusheng Xu, Wenjie Ye, Weilin Liu 외

Reward models have been increasingly critical for improving the reasoning capability of LLMs. Existing research has shown that a well-trained reward model can substantially improve model performances at inference time vi…

GSM8KMath

Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards

2025-05-07 · Yuxin Zhang, Meihao Fan, Ju Fan, Mingyang Yi 외

Recent advances in large language models (LLMs) have significantly improved performance on the Text-to-SQL task by leveraging their powerful reasoning capabilities. To enhance accuracy during the reasoning process, exter…

Text to SQLText-To-SQL

Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment

2024-02-15 · Rui Yang, Xiaoman Pan, Feng Luo, Shuang Qiu 외

We consider the problem of multi-objective alignment of foundation models with human preferences, which is a critical step towards helpful and harmless AI systems. However, it is generally costly and unstable to fine-tun…

GPUReinforcement Learning (RL)

Safe Imitation Learning via Fast Bayesian Reward Inference from Preferences

2020-02-21 · ICML 2020 1 · Daniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott Niekum

Bayesian reward learning from demonstrations enables rigorous safety and uncertainty analysis when performing imitation learning. However, Bayesian reward learning methods are typically computationally intractable for co…

Atari GamesBayesian InferenceImitation Learning