paper-with-me

홈 › Papers

Knowledge is reward: Learning optimal exploration by predictive reward cashing

2021-09-17 · Luca Ambrogioni

There is a strong link between the general concept of intelligence and the ability to collect and use information. The theory of Bayes-adaptive exploration offers an attractive optimality framework for training machines to perform complex information gathering tasks. However, the computational complexity of the resulting optimal control problem has limited the diffusion of the theory to mainstream deep AI research. In this paper we exploit the inherent mathematical structure of Bayes-adaptive problems in order to dramatically simplify the problem by making the reward structure denser while simultaneously decoupling the learning of exploitation and exploration policies. The key to this simplification comes from the novel concept of cross-value (i.e. the value of being in an environment while acting optimally according to another), which we use to quantify the value of currently available information. This results in a new denser reward structure that "cashes in" all future rewards that can be predicted from the current information state. In a set of experiments we show that the approach makes it possible to learn challenging information gathering tasks without the use of shaping and heuristic bonuses in situations where the standard RL algorithms fail.

📄 PDF Abstract BibTeX arXiv:2109.08518

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Exploration Unbound

2024-07-16 · Dilip Arumugam, Wanqiao Xu, Benjamin Van Roy

A sequential decision-making agent balances between exploring to gain new knowledge about an environment and exploiting current knowledge to maximize immediate reward. For environments studied in the traditional literatu…

Decision MakingSequential Decision Making

Improved Bounds for Reward-Agnostic and Reward-Free Exploration

2026-02-18 · Oran Ridel, Alon Cohen arxiv

We study reward-free and reward-agnostic exploration in episodic finite-horizon Markov decision processes (MDPs), where an agent explores an unknown environment without observing external rewards. Reward-free exploration…

Reward-Free RL is No Harder Than Reward-Aware RL in Linear Markov Decision Processes

2022-01-26 · Andrew Wagenmaker, Yifang Chen, Max Simchowitz, Simon S. Du 외

Reward-free reinforcement learning (RL) considers the setting where the agent does not have access to a reward function during exploration, but must propose a near-optimal policy for an arbitrary reward function revealed…

Reinforcement Learning (RL)

Prioritized Guidance for Efficient Multi-Agent Reinforcement Learning Exploration

2019-07-18 · Qisheng Wang, Qichao Wang

Exploration efficiency is a challenging problem in multi-agent reinforcement learning (MARL), as the policy learned by confederate MARL depends on the collaborative approach among multiple agents. Another important probl…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Active Finite Reward Automaton Inference and Reinforcement Learning Using Queries and Counterexamples

2020-06-28 · Zhe Xu, Bo Wu, Aditya Ojha, Daniel Neider 외

Despite the fact that deep reinforcement learning (RL) has surpassed human-level performances in various tasks, it still has several fundamental challenges. First, most RL methods require intensive data from the explorat…

Active LearningDeep Reinforcement LearningQ-Learningreinforcement-learning+1