paper-with-me

홈 › Papers

Pavlovian Signalling with General Value Functions in Agent-Agent Temporal Decision Making

2022-01-11 · Andrew Butcher, Michael Bradley Johanson, Elnaz Davoodi, Dylan J. A. Brenneis, Leslie Acker, Adam S. R. Parker, Adam White, Joseph Modayil, Patrick M. Pilarski

In this paper, we contribute a multi-faceted study into Pavlovian signalling -- a process by which learned, temporally extended predictions made by one agent inform decision-making by another agent. Signalling is intimately connected to time and timing. In service of generating and receiving signals, humans and other animals are known to represent time, determine time since past events, predict the time until a future stimulus, and both recognize and generate patterns that unfold in time. We investigate how different temporal processes impact coordination and signalling between learning agents by introducing a partially observable decision-making domain we call the Frost Hollow. In this domain, a prediction learning agent and a reinforcement learning agent are coupled into a two-part decision-making system that works to acquire sparse reward while avoiding time-conditional hazards. We evaluate two domain variations: machine agents interacting in a seven-state linear walk, and human-machine interaction in a virtual-reality environment. Our results showcase the speed of learning for Pavlovian signalling, the impact that different temporal representations do (and do not) have on agent-agent coordination, and how temporal aliasing impacts agent-agent and human-agent interactions differently. As a main contribution, we establish Pavlovian signalling as a natural bridge between fixed signalling paradigms and fully adaptive communication learning between two agents. We further show how to computationally build this adaptive signalling process out of a fixed signalling process, characterized by fast continual prediction learning and minimal constraints on the nature of the agent receiving signals. Our results therefore suggest an actionable, constructivist path towards communication learning between reinforcement learning agents.

📄 PDF Abstract BibTeX arXiv:2201.03709

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Makingreinforcement-learningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

The Frost Hollow Experiments: Pavlovian Signalling as a Path to Coordination and Communication Between Agents

2022-03-17 · Patrick M. Pilarski, Andrew Butcher, Elnaz Davoodi, Michael Bradley Johanson 외

Learned communication between agents is a powerful tool when approaching decision-making problems that are hard to overcome by any single agent in isolation. However, continual coordination and communication learning bet…

Decision Makingreinforcement-learningReinforcement Learning (RL)

Continually Learned Pavlovian Signalling Without Forgetting for Human-in-the-Loop Robotic Control

2023-05-16 · Adam S. R. Parker, Michael R. Dawson, Patrick M. Pilarski

Artificial limbs are sophisticated devices to assist people with tasks of daily living. Despite advanced robotic prostheses demonstrating similar motion capabilities to biological limbs, users report them difficult and n…

Mathematical Modelling and Analysis of the Brassinosteroid and Gibberellin Signalling Pathways and their Interactions

2017-08-19

The plant hormones brassinosteroid (BR) and gibberellin (GA) have important roles in a wide range of processes involved in plant growth and development. In this paper we derive and analyse new mathematical models for the…

Monadic Pavlovian associative learning in a backpropagation-free photonic network

2020-11-30 · James Y. S. Tan, Zengguang Cheng, Johannes Feldmann, Xuan Li 외

Over a century ago, Ivan P. Pavlov, in a classic experiment, demonstrated how dogs can learn to associate a ringing bell with food, thereby causing a ring to result in salivation. Today, it is rare to find the use of Pav…

Good signals gone bad: dynamic signalling with switching efforts

2017-07-15

This paper examines signalling when the sender exerts effort and receives benefits over time. Receivers only observe a noisy public signal about the effort, which has no intrinsic value. The modelling of signalling in …

Vocal Bursts Type Prediction