AUPO -- Abstracted Until Proven Otherwise: A Reward Distribution Based Abstraction Algorithm
We introduce a novel, drop-in modification to Monte Carlo Tree Search's (MCTS) decision policy that we call AUPO. Comparisons based on a range of IPPC benchmark problems show that AUPO clearly outperforms MCTS. AUPO is an automatic action abstraction algorithm that solely relies on reward distribution statistics acquired during the MCTS. Thus, unlike other automatic abstraction algorithms, AUPO requires neither access to transition probabilities nor does AUPO require a directed acyclic search graph to build its abstraction, allowing AUPO to detect symmetric actions that state-of-the-art frameworks like ASAP struggle with when the resulting symmetric states are far apart in state space. Furthermore, as AUPO only affects the decision policy, it is not mutually exclusive with other abstraction techniques that only affect the tree search.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Learning Soft Constraints From Constrained Expert Demonstrations
Inverse reinforcement learning (IRL) methods assume that the expert data is generated by an agent optimizing some reward function. However, in many settings, the agent may optimize a reward function subject to some const…
Attention-Based Reward Shaping for Sparse and Delayed Rewards
Sparse and delayed reward functions pose a significant obstacle for real-world Reinforcement Learning (RL) applications. In this work, we propose Attention-based REward Shaping (ARES), a general and robust algorithm whic…
Reinforcement Learning (RL)Microbial Mat Metagenomes from Waikite Valley, Aotearoa New Zealand
The rise of complex multicellular ecosystems Neoproterozoic time was preceded by a microbial Proterozoic biosphere, where productivity may have been largely restricted to microbial mats made up of bacteria including oxyg…
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
We study online preference-based reinforcement learning (PbRL) with the goal of improving sample efficiency. While a growing body of theoretical work has emerged-motivated by PbRL's recent empirical success, particularly…
Reinforcement LearningA Closer Look at Reward Decomposition for High-Level Robotic Explanations
Explaining the behaviour of intelligent agents learned by reinforcement learning (RL) to humans is challenging yet crucial due to their incomprehensible proprioceptive states, variational intermediate goals, and resultan…
Reinforcement Learning (RL)Vocal Bursts Intensity Prediction