paper-with-me

홈 › Papers

Hallucinating Value: A Pitfall of Dyna-style Planning with Imperfect Environment Models

2020-06-08 · Taher Jafferjee, Ehsan Imani, Erin Talvitie, Martha White, Micheal Bowling

Dyna-style reinforcement learning (RL) agents improve sample efficiency over model-free RL agents by updating the value function with simulated experience generated by an environment model. However, it is often difficult to learn accurate models of environment dynamics, and even small errors may result in failure of Dyna agents. In this paper, we investigate one type of model error: hallucinated states. These are states generated by the model, but that are not real states of the environment. We present the Hallucinated Value Hypothesis (HVH): updating values of real states towards values of hallucinated states results in misleading state-action values which adversely affect the control policy. We discuss and evaluate four Dyna variants; three which update real states toward simulated -- and therefore potentially hallucinated -- states and one which does not. The experimental results provide evidence for the HVH thus suggesting a fruitful direction toward developing Dyna algorithms robust to model error.

📄 PDF Abstract BibTeX arXiv:2006.04363

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping

2012-06-13 · Richard S. Sutton, Csaba Szepesvari, Alborz Geramifard, Michael P. Bowling

We consider the problem of efficiently learning optimal control policies and value functions over large state spaces in an online setting in which estimates must be available after each interaction with the world. This p…

Learning from Hallucinating Critical Points for Navigation in Dynamic Environments

2025-09-30 · Saad Abdul Ghani, Kameron Lee, Xuesu Xiao arxiv

Generating large and diverse obstacle datasets to learn motion planning in environments with dynamic obstacles is challenging due to the vast space of possible obstacle trajectories. Inspired by hallucination-based data …

Motion Planning

Uncertainty - sensitive learning and planning with ensembles

2019-09-25 · Piotr Miłoś, Łukasz Kuciński, Konrad Czechowski, Piotr Kozakowski 외

We propose a reinforcement learning framework for discrete environments in which an agent optimizes its behavior on two timescales. For the short one, it uses tree search methods to perform tactical decisions. The long s…

Montezuma's RevengeSokoban

Hallucinating Agnostic Images to Generalize Across Domains

2018-08-03 · Fabio M. Carlucci, Paolo Russo, Tatiana Tommasi, Barbara Caputo

The ability to generalize across visual domains is crucial for the robustness of artificial recognition systems. Although many training sources may be available in real contexts, the access to even unlabeled target sampl…

Domain AdaptationDomain GeneralizationUnsupervised Domain Adaptation

Can a Hallucinating Model help in Reducing Human "Hallucination"?

2024-05-01 · Sowmya S Sundaram, Balaji Alwar

The prevalence of unwarranted beliefs, spanning pseudoscience, logical fallacies, and conspiracy theories, presents substantial societal hurdles and the risk of disseminating misinformation. Utilizing established psychom…

HallucinationLogical FallaciesMisconceptionsMisinformation