Deep Learning Agents Trained For Avoidance Behave Like Hawks And Doves
We present heuristically optimal strategies expressed by deep learning agents playing a simple avoidance game. We analyse the learning and behaviour of two agents within a symmetrical grid world that must cross paths to reach a target destination without crashing into each other or straying off of the grid world in the wrong direction. The agent policy is determined by one neural network that is employed in both agents. Our findings indicate that the fully trained network exhibits behaviour similar to that of the game Hawks and Doves, in that one agent employs an aggressive strategy to reach the target while the other learns how to avoid the aggressive agent.
Code (1)
Similar Papers 제목 키워드 기반
Multi-agent navigation based on deep reinforcement learning and traditional pathfinding algorithm
We develop a new framework for multi-agent collision avoidance problem. The framework combined traditional pathfinding algorithm and reinforcement learning. In our approach, the agents learn whether to be navigated or to…
Collision AvoidanceDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1Human-Inspired Multi-Agent Navigation using Knowledge Distillation
Despite significant advancements in the field of multi-agent navigation, agents still lack the sophistication and intelligence that humans exhibit in multi-agent settings. In this paper, we propose a framework for learni…
Collision AvoidanceKnowledge Distillationreinforcement-learningReinforcement Learning (RL)HyperSpacetime: Complex Algebro-Geometric Analysis of Intelligence Quantum Entanglement Convergent Evolution (Extended Abstract)
Far from the Madding Cloud When groups of ants in haystack as agents Descending gudied by curvatures instead of gradients No powerful enough in changing their environments Being inferior adapters instead of equal game…
Meta-trained agents implement Bayes-optimal agents
Memory-based meta-learning is a powerful technique to build agents that adapt fast to any task within a target distribution. A previous theoretical study has argued that this remarkable performance is because the meta-tr…
Meta-LearningAn Interpretable Data-Driven Model of the Flight Dynamics of Hawks
Despite significant analysis of bird flight, generative physics models for flight dynamics do not currently exist. Yet the underlying mechanisms responsible for various flight manoeuvres are important for understanding h…