Transfer of Fully Convolutional Policy-Value Networks Between Games and Game Variants
In this paper, we use fully convolutional architectures in AlphaZero-like self-play training setups to facilitate transfer between variants of board games as well as distinct games. We explore how to transfer trained parameters of these architectures based on shared semantics of channels in the state and action representations of the Ludii general game system. We use Ludii's large library of games and game variants for extensive transfer learning evaluations, in zero-shot transfer experiments as well as experiments with additional fine-tuning time.
Code (0)
등록된 구현이 없습니다.
Tasks
Board GamesTransfer LearningSimilar Papers 제목 키워드 기반
Successor Features for Transfer in Reinforcement Learning
Transfer in reinforcement learning refers to the notion that generalization should occur not only within a task but also across tasks. We propose a transfer framework for the scenario where the reward function changes be…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)CURO: Curriculum Learning for Relative Overgeneralization
Relative overgeneralization (RO) is a pathology that can arise in cooperative multi-agent tasks when the optimal joint action's utility falls below that of a sub-optimal joint action. RO can cause the agents to get stuck…
Efficient ExplorationMulti-agent Reinforcement LearningStarcraftStarcraft II+1RL-Driven Sustainable Land-Use Allocation for the Lake Malawi Basin
Unsustainable land-use practices in ecologically sensitive regions threaten biodiversity, water resources, and the livelihoods of millions. This paper presents a deep reinforcement learning (RL) framework for optimizing …
Reinforcement LearningMulti-agent Policy Reciprocity with Theoretical Guarantee
Modern multi-agent reinforcement learning (RL) algorithms hold great potential for solving a variety of real-world problems. However, they do not fully exploit cross-agent knowledge to reduce sample complexity and improv…
continuous-controlContinuous ControlMulti-agent Reinforcement LearningReinforcement Learning (RL)Soft Value Iteration Networks for Planetary Rover Path Planning
Value iteration networks are an approximation of the value iteration (VI) algorithm implemented with convolutional neural networks to make VI fully differentiable. In this work, we study these networks in the context of …
Motion Planning