Eventual Discounting Temporal Logic Counterfactual Experience Replay
Linear temporal logic (LTL) offers a simplified way of specifying tasks for policy optimization that may otherwise be difficult to describe with scalar reward functions. However, the standard RL framework can be too myopic to find maximally LTL satisfying policies. This paper makes two contributions. First, we develop a new value-function based proxy, using a technique we call eventual discounting, under which one can find policies that satisfy the LTL specification with highest achievable probability. Second, we develop a new experience replay method for generating off-policy data from on-policy rollouts via counterfactual reasoning on different ways of satisfying the LTL specification. Our experiments, conducted in both discrete and continuous state-action spaces, confirm the effectiveness of our counterfactual experience replay approach.
Code (0)
등록된 구현이 없습니다.
Tasks
counterfactualCounterfactual ReasoningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Discounting and Drug Seeking in Biological Hierarchical Reinforcement Learning
Despite a strong desire to quit, individuals with long-term substance use disorder (SUD) often struggle to resist drug use, even when aware of its harmful consequences. This disconnect between knowledge and compulsive be…
Hierarchical Reinforcement Learningreinforcement-learningReinforcement LearningDecision-making and Fuzzy Temporal Logic
This paper shows that the fuzzy temporal logic can model figures of thought to describe decision-making behaviors. In order to exemplify, some economic behaviors observed experimentally were modeled from problems of choi…
Decision MakingConservative, Proportional and Optimistic Contextual Discounting in the Belief Functions Theory
Information discounting plays an important role in the theory of belief functions and, generally, in information fusion. Nevertheless, neither classical uniform discounting nor contextual cannot model certain use cases, …
The role of time estimation in decreased impatience in Intertemporal Choice
The role of specific cognitive processes in deviations from constant discounting in intertemporal choice is not well understood. We evaluated decreased impatience in intertemporal choice tasks independent of discounting …
The Problem of Coincidence in A Theory of Temporal Multiple Recurrence
Logical theories have been developed which have allowed temporal reasoning about eventualities (a la Galton) such as states, processes, actions, events, processes and complex eventualities such as sequences and recurrenc…