paper-with-me

홈 › Papers

AUPO -- Abstracted Until Proven Otherwise: A Reward Distribution Based Abstraction Algorithm

2025-10-27 · Robin Schmöcker, Alexander Dockhorn, Bodo Rosenhahn arxiv

We introduce a novel, drop-in modification to Monte Carlo Tree Search's (MCTS) decision policy that we call AUPO. Comparisons based on a range of IPPC benchmark problems show that AUPO clearly outperforms MCTS. AUPO is an automatic action abstraction algorithm that solely relies on reward distribution statistics acquired during the MCTS. Thus, unlike other automatic abstraction algorithms, AUPO requires neither access to transition probabilities nor does AUPO require a directed acyclic search graph to build its abstraction, allowing AUPO to detect symmetric actions that state-of-the-art frameworks like ASAP struggle with when the resulting symmetric states are far apart in state space. Furthermore, as AUPO only affects the decision policy, it is not mutually exclusive with other abstraction techniques that only affect the tree search.

📄 PDF Abstract BibTeX arXiv:2510.23214

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Soft Constraints From Constrained Expert Demonstrations

2022-06-02 · Ashish Gaurav, Kasra Rezaee, Guiliang Liu, Pascal Poupart

Inverse reinforcement learning (IRL) methods assume that the expert data is generated by an agent optimizing some reward function. However, in many settings, the agent may optimize a reward function subject to some const…

Attention-Based Reward Shaping for Sparse and Delayed Rewards

2025-05-16 · Ian Holmes, Min Chi

Sparse and delayed reward functions pose a significant obstacle for real-world Reinforcement Learning (RL) applications. In this work, we propose Attention-based REward Shaping (ARES), a general and robust algorithm whic…

Reinforcement Learning (RL)

Microbial Mat Metagenomes from Waikite Valley, Aotearoa New Zealand

2024-12-02 · Beatrice Tauer, Elizabeth Trembath-Reichert, L. M. Ward

The rise of complex multicellular ecosystems Neoproterozoic time was preceded by a microbial Proterozoic biosphere, where productivity may have been largely restricted to microbial mats made up of bacteria including oxyg…

Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options

2025-10-21 · Joongkyu Lee, Seouh-won Yi, Min-hwan Oh arxiv

We study online preference-based reinforcement learning (PbRL) with the goal of improving sample efficiency. While a growing body of theoretical work has emerged-motivated by PbRL's recent empirical success, particularly…

Reinforcement Learning

A Closer Look at Reward Decomposition for High-Level Robotic Explanations

2023-04-25 · Wenhao Lu, Xufeng Zhao, Sven Magg, Martin Gromniak 외

Explaining the behaviour of intelligent agents learned by reinforcement learning (RL) to humans is challenging yet crucial due to their incomprehensible proprioceptive states, variational intermediate goals, and resultan…

Reinforcement Learning (RL)Vocal Bursts Intensity Prediction