ToMCAT: Theory-of-Mind for Cooperative Agents in Teams via Multiagent Diffusion Policies
In this paper we present ToMCAT (Theory-of-Mind for Cooperative Agents in Teams), a new framework for generating ToM-conditioned trajectories. It combines a meta-learning mechanism, that performs ToM reasoning over teammates' underlying goals and future behavior, with a multiagent denoising-diffusion model, that generates plans for an agent and its teammates conditioned on both the agent's goals and its teammates' characteristics, as computed via ToM. We implemented an online planning system that dynamically samples new trajectories (replans) from the diffusion model whenever it detects a divergence between a previously generated plan and the current state of the world. We conducted several experiments using ToMCAT in a simulated cooking domain. Our results highlight the importance of the dynamic replanning mechanism in reducing the usage of resources without sacrificing team performance. We also show that recent observations about the world and teammates' behavior collected by an agent over the course of an episode combined with ToM inferences are crucial to generate team-aware plans for dynamic adaptation to teammates, especially when no prior information is provided about them.
Code (0)
등록된 구현이 없습니다.
Tasks
DenoisingMeta-LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Improving Multi-Agent Cooperation using Theory of Mind
Recent advances in Artificial Intelligence have produced agents that can beat human world champions at games like Go, Starcraft, and Dota2. However, most of these models do not seem to play in a human-like manner: People…
StarcraftThe ToMCAT Dataset
We present a rich, multimodal dataset consisting of data from 40 teams of three humans conducting simulated urban search-and-rescue (SAR) missions in a Minecraft-based testbed, collected for the Theory of Mind-based Cogn…
Theory of Mind with Guilt Aversion Facilitates Cooperative Reinforcement Learning
Guilt aversion induces experience of a utility loss in people if they believe they have disappointed others, and this promotes cooperative behaviour in human. In psychological game theory, guilt aversion necessitates mod…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Theory of Mind for Deep Reinforcement Learning in Hanabi
The partially observable card game Hanabi has recently been proposed as a new AI challenge problem due to its dependence on implicit communication conventions and apparent necessity of theory of mind reasoning for effici…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Probabilistic Modeling of Human Teams to Infer False Beliefs
We develop a probabilistic graphical model (PGM) for artificially intelligent (AI) agents to infer human beliefs during a simulated urban search and rescue (USAR) scenario executed in a Minecraft environment with a team …
AI AgentMinecraft