paper-with-me

홈 › Papers

Optimal Behavior Prior: Data-Efficient Human Models for Improved Human-AI Collaboration

2022-11-03 · Mesut Yang, Micah Carroll, Anca Dragan

AI agents designed to collaborate with people benefit from models that enable them to anticipate human behavior. However, realistic models tend to require vast amounts of human data, which is often hard to collect. A good prior or initialization could make for more data-efficient training, but what makes for a good prior on human behavior? Our work leverages a very simple assumption: people generally act closer to optimal than to random chance. We show that using optimal behavior as a prior for human models makes these models vastly more data-efficient and able to generalize to new environments. Our intuition is that such a prior enables the training to focus one's precious real-world data on capturing the subtle nuances of human suboptimality, instead of on the basics of how to do the task in the first place. We also show that using these improved human models often leads to better human-AI collaboration performance compared to using models based on real human data alone.

📄 PDF Abstract BibTeX arXiv:2211.01602

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Human-Variability-Respecting Optimal Control for Physical Human-Machine Interaction

2024-05-06 · Sean Kille, Paul Leibold, Philipp Karg, Balint Varga 외

Physical Human-Machine Interaction plays a pivotal role in facilitating collaboration across various domains. When designing appropriate model-based controllers to assist a human in the interaction, the accuracy of the h…

The Boltzmann Policy Distribution: Accounting for Systematic Suboptimality in Human Models

2022-04-22 · ICLR 2022 4 · Cassidy Laidlaw, Anca Dragan

Models of human behavior for prediction and collaboration tend to fall into two categories: ones that learn from large amounts of data via imitation learning, and ones that assume human behavior to be noisily-optimal for…

Bayesian InferenceImitation Learning

On Gap-dependent Bounds for Offline Reinforcement Learning

2022-06-01 · Xinqi Wang, Qiwen Cui, Simon S. Du

This paper presents a systematic study on gap-dependent sample complexity in offline reinforcement learning. Prior work showed when the density ratio between an optimal policy and the behavior policy is upper bounded (th…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Enhancing Goal Inference via Correction Timing

2026-02-20 · Anjiabei Wang, Shuangge Wang, Tesca Fitzgerald arxiv

Corrections offer a natural modality for people to provide feedback to a robot, by (i) intervening in the robot's behavior when they believe the robot is failing (or will fail) the task objectives and (ii) modifying the …

Bayesian Optimal Experimental Design for Simulator Models of Cognition

2021-10-29 · NeurIPS Workshop AI4Scien 2021 12 · Simon Valentin, Steven Kleinegesse, Neil R. Bramley, Michael U. Gutmann 외

Bayesian optimal experimental design (BOED) is a methodology to identify experiments that are expected to yield informative data. Recent work in cognitive science considered BOED for computational models of human behavio…

Experimental Designparameter estimation