The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior
In this paper we illustrate a novel architecture generating interpretable behavior and explanations. We refer to this architecture as the Mirror Agent Model because it defines the observer model, that is the target of explicit and implicit communications, as a mirror of the agent's. With the goal of providing a general understanding of this work, we firstly show prior relevant results addressing the informative communication of agents intentions and the production of legible behavior. In the second part of the paper we furnish the architecture with novel capabilities for explanations through off-the-shelf saliency methods, followed by preliminary qualitative results.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Toward Universal and Interpretable World Models for Open-ended Learning Agents
We introduce a generic, compositional and interpretable class of generative world models that supports open-ended learning agents. This is a sparse class of Bayesian networks capable of approximating a broad range of sto…
Developmental LearningDICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination
Multi-agent large language model (LLM) systems often fail to reliably outperform a single strong model equipped with best-of-N sampling. We argue that a core source of this instability is ill-posed equilibrium selection:…
Mirror: A Multi-Agent System for AI-Assisted Ethics Review
Ethics review is a foundational mechanism of modern research governance, yet contemporary systems face increasing strain as ethical risks arise as structural consequences of large-scale, interdisciplinary scientific prac…
Distributed Learning with Infinitely Many Hypotheses
We consider a distributed learning setup where a network of agents sequentially access realizations of a set of random variables with unknown distributions. The network objective is to find a parametrized distribution th…
Distributed Learning for Cooperative Inference
We study the problem of cooperative inference where a group of agents interact over a network and seek to estimate a joint parameter that best explains a set of observations. Agents do not know the network topology or th…