paper-with-me

홈 › Papers

The Frost Hollow Experiments: Pavlovian Signalling as a Path to Coordination and Communication Between Agents

2022-03-17 · Patrick M. Pilarski, Andrew Butcher, Elnaz Davoodi, Michael Bradley Johanson, Dylan J. A. Brenneis, Adam S. R. Parker, Leslie Acker, Matthew M. Botvinick, Joseph Modayil, Adam White

Learned communication between agents is a powerful tool when approaching decision-making problems that are hard to overcome by any single agent in isolation. However, continual coordination and communication learning between machine agents or human-machine partnerships remains a challenging open problem. As a stepping stone toward solving the continual communication learning problem, in this paper we contribute a multi-faceted study into what we term Pavlovian signalling -- a process by which learned, temporally extended predictions made by one agent inform decision-making by another agent with different perceptual access to their shared environment. We seek to establish how different temporal processes and representational choices impact Pavlovian signalling between learning agents. To do so, we introduce a partially observable decision-making domain we call the Frost Hollow. In this domain a prediction learning agent and a reinforcement learning agent are coupled into a two-part decision-making system that seeks to acquire sparse reward while avoiding time-conditional hazards. We evaluate two domain variations: 1) machine prediction and control learning in a linear walk, and 2) a prediction learning machine interacting with a human participant in a virtual reality environment. Our results showcase the speed of learning for Pavlovian signalling, the impact that different temporal representations do (and do not) have on agent-agent coordination, and how temporal aliasing impacts agent-agent and human-agent interactions differently. As a main contribution, we establish Pavlovian signalling as a natural bridge between fixed signalling paradigms and fully adaptive communication learning. Our results therefore point to an actionable, constructivist path towards continual communication learning between reinforcement learning agents, with potential impact in a range of real-world settings.

📄 PDF Abstract BibTeX arXiv:2203.09498

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Makingreinforcement-learningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Pavlovian Signalling with General Value Functions in Agent-Agent Temporal Decision Making

2022-01-11 · Andrew Butcher, Michael Bradley Johanson, Elnaz Davoodi, Dylan J. A. Brenneis 외

In this paper, we contribute a multi-faceted study into Pavlovian signalling -- a process by which learned, temporally extended predictions made by one agent inform decision-making by another agent. Signalling is intimat…

Decision Makingreinforcement-learningReinforcement Learning (RL)

Continually Learned Pavlovian Signalling Without Forgetting for Human-in-the-Loop Robotic Control

2023-05-16 · Adam S. R. Parker, Michael R. Dawson, Patrick M. Pilarski

Artificial limbs are sophisticated devices to assist people with tasks of daily living. Despite advanced robotic prostheses demonstrating similar motion capabilities to biological limbs, users report them difficult and n…

Mathematical Modelling and Analysis of the Brassinosteroid and Gibberellin Signalling Pathways and their Interactions

2017-08-19

The plant hormones brassinosteroid (BR) and gibberellin (GA) have important roles in a wide range of processes involved in plant growth and development. In this paper we derive and analyse new mathematical models for the…

FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning

2026-01-26 · Haozheng Luo, Zhuolin Jiang, Md Zahid Hasan, Yan Chen 외 arxiv

We propose FROST, an attention-aware method for efficient reasoning. Unlike traditional approaches, FROST leverages attention weights to prune uncritical reasoning paths, yielding shorter and more reliable reasoning traj…

Cellular forgetting, desensitisation, stress and aging in signalling networks. When do cells refuse to learn more?

2023-12-28 · Tamas Veres, Mark Kerestely, Borbala M. Kovacs, David Keresztes 외

Recent findings show that single, non-neuronal cells are also able to learn signalling responses developing cellular memory. In cellular learning nodes of signalling networks strengthen their interactions e.g. by the con…

Change DetectionDrug Design