paper-with-me

홈 › Papers

How Much Progress Did I Make? An Unexplored Human Feedback Signal for Teaching Robots

2024-07-08 · Hang Yu, Qidi Fang, Shijie Fang, Reuben M. Aronson, Elaine Schaertl Short

Enhancing the expressiveness of human teaching is vital for both improving robots' learning from humans and the human-teaching-robot experience. In this work, we characterize and test a little-used teaching signal: \textit{progress}, designed to represent the completion percentage of a task. We conducted two online studies with 76 crowd-sourced participants and one public space study with 40 non-expert participants to validate the capability of this progress signal. We find that progress indicates whether the task is successfully performed, reflects the degree of task completion, identifies unproductive but harmless behaviors, and is likely to be more consistent across participants. Furthermore, our results show that giving progress does not require extra workload and time. An additional contribution of our work is a dataset of 40 non-expert demonstrations from the public space study through an ice cream topping-adding task, which we observe to be multi-policy and sub-optimal, with sub-optimality not only from teleoperation errors but also from exploratory actions and attempts. The dataset is available at \url{https://github.com/TeachingwithProgress/Non-Expert\_Demonstrations}.

📄 PDF Abstract BibTeX arXiv:2407.06459

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Selective Progress-Aware Querying for Human-in-the-Loop Reinforcement Learning

2025-09-24 · Anujith Muraleedharan, Anamika J H arxiv

Human feedback can greatly accelerate robot learning, but in real-world settings, such feedback is costly and limited. Existing human-in-the-loop reinforcement learning (HiL-RL) methods often assume abundant feedback, li…

Reinforcement Learning

Let Me Teach You: Pedagogical Foundations of Feedback for Language Models

2023-07-01 · Beatriz Borges, Niket Tandon, Tanja Käser, Antoine Bosselut

Natural Language Feedback (NLF) is an increasingly popular mechanism for aligning Large Language Models (LLMs) to human preferences. Despite the diversity of the information it can convey, NLF methods are often hand-desi…

Diversity

Stochastic contextual bandits with graph feedback: from independence number to MAS number

2024-02-12 · Yuxiao Wen, Yanjun Han, Zhengyuan Zhou

We consider contextual bandits with graph feedback, a class of interactive learning problems with richer structures than vanilla contextual bandits, where taking an action reveals the rewards for all neighboring actions …

Multi-Armed Bandits

VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

2024-06-21 · Xuan He, Dongfu Jiang, Ge Zhang, Max Ku 외

The recent years have witnessed great advances in video generation. However, the development of automatic video metrics is lagging significantly behind. None of the existing metric is able to provide reliable scores over…

Video GenerationVideo Quality Assessment

Exploiting Unlabeled Data for Feedback Efficient Human Preference based Reinforcement Learning

2023-02-17 · Mudit Verma, Siddhant Bhambri, Subbarao Kambhampati

Preference Based Reinforcement Learning has shown much promise for utilizing human binary feedback on queried trajectory pairs to recover the underlying reward model of the Human in the Loop (HiL). While works have attem…

reinforcement-learningReinforcement Learning (RL)