paper-with-me

홈 › Papers

Variational oracle guiding for reinforcement learning

2021-09-29 · ICLR 2022 4 · Dongqi Han, Tadashi Kozuno, Xufang Luo, Zhao-Yun Chen, Kenji Doya, Yuqing Yang, Dongsheng Li

How to make intelligent decisions is a central problem in machine learning and cognitive science. Despite recent successes of deep reinforcement learning (RL) in various decision making problems, an important but under-explored aspect is how to leverage oracle observation (the information that is invisible during online decision making, but is available during offline training) to facilitate learning. For example, human experts will look at the replay after a Poker game, in which they can check the opponents' hands to improve their estimation of the opponents' hands from the visible information during playing. In this work, we study such problems based on Bayesian theory and derive an objective to leverage oracle observation in RL using variational method. Our key contribution is to propose a general learning framework referred to as variational latent oracle guiding (VLOG) for deep RL. VLOG is featured with preferable properties such as its robust and promising performance and its versatility to incorporate with any value-based deep RL algorithm. We empirically demonstrate the effectiveness of VLOG in online and offline RL domains using decision-making tasks ranged from video games to a challenging tile-based game Mahjong. Furthermore, we publish the environment of Mahjong and the corresponding offline RL dataset as a benchmark to facilitate future research on oracle guiding.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingDeep Reinforcement LearningOffline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Oracle-guided Dynamic User Preference Modeling for Sequential Recommendation

2024-12-01 · Jiafeng Xia, Dongsheng Li, Hansu Gu, Tun Lu 외

Sequential recommendation methods can capture dynamic user preferences from user historical interactions to achieve better performance. However, most existing methods only use past information extracted from user histori…

Sequential Recommendation

Reference-free logged energy-oracle recovery for neural approximations of symmetric coercive variational problems: conforming Riesz reconstruction and archive-level selection

2026-08-17 · Karim Bounja, Lahcen Laayouni, Boujemaa Achchab, Abdeljalil Sakat arxiv

Neural PDE training yields a finite checkpoint archive, yet its logged energy errors are inaccessible without the exact solution, while loss-based selection does not necessarily recover the logged energy oracle. For admi…

Approximate Dynamic Oracle for Dependency Parsing with Reinforcement Learning

2018-11-01 · WS 2018 11 · Xiang Yu, Ngoc Thang Vu, Jonas Kuhn

We present a general approach with reinforcement learning (RL) to approximate dynamic oracles for transition systems where exact dynamic oracles are difficult to derive. We treat oracle parsing as a reinforcement learnin…

Dependency ParsingImitation LearningQ-Learningreinforcement-learning+3

Opinion-Guided Reinforcement Learning

2024-05-27 · Kyanna Dagenais, Istvan David

Human guidance is often desired in reinforcement learning to improve the performance of the learning agent. However, human insights are often mere opinions and educated guesses rather than well-formulated arguments. Whil…

Efficient Explorationreinforcement-learningReinforcement Learning

Oracle Supervision Transfers for Hyperparameter Prediction in Model-Based Image Denoising

2026-05-19 · Jianmin Liao, Lixin Shen, Yuesheng Xu arxiv

Hyperparameter prediction is a critical practical bottleneck for model-based image denoisers, ranging from classical TV/TGV variational solvers to modern diffusion-based models such as DiffPIR. While existing learned pre…

Image Denoising