Generalizing LTL Instructions via Future Dependent Options
In many real-world applications of control system and robotics, linear temporal logic (LTL) is a widely-used task specification language which has a compositional grammar that naturally induces temporally extended behaviours across tasks, including conditionals and alternative realizations. An important problem in RL with LTL tasks is to learn task-conditioned policies which can zero-shot generalize to new LTL instructions not observed in the training. However, because symbolic observation is often lossy and LTL tasks can have long time horizon, previous works can suffer from issues such as training sampling inefficiency and infeasibility or sub-optimality of the found solutions. In order to tackle these issues, this paper proposes a novel multi-task RL algorithm with improved learning efficiency and optimality. To achieve the global optimality of task completion, we propose to learn options dependent on the future subgoals via a novel off-policy approach. In order to propagate the rewards of satisfying future subgoals back more efficiently, we propose to train a multi-step value function conditioned on the subgoal sequence which is updated with Monte Carlo estimates of multi-step discounted returns. In experiments on three different domains, we evaluate the LTL generalization capability of the agent trained by the proposed method, showing its advantage over previous representative methods.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Semi-analytical pricing of options written on SOFR futures
In this paper, we propose a semi-analytical approach to pricing options on SOFR futures where the underlying SOFR follows a time-dependent CEV model. By definition, these options change their type at the beginning of the…
Evaluating the Instruction-following Abilities of Language Models using Knowledge Tasks
In this work, we focus our attention on developing a benchmark for instruction-following where it is easy to verify both task performance as well as instruction-following capabilities. We adapt existing knowledge benchma…
Instruction FollowingMultiple-choiceRobust Trading of Implied Skew
In this paper, we present a method for constructing a (static) portfolio of co-maturing European options whose price sign is determined by the skewness level of the associated implied volatility. This property holds rega…
From the Samuelson Volatility Effect to a Samuelson Correlation Effect: Evidence from Crude Oil Calendar Spread Options
We introduce a multi-factor stochastic volatility model based on the CIR/Heston stochastic volatility process. In order to capture the Samuelson effect displayed by commodity futures contracts, we add expiry-dependent ex…
American options valuation in time-dependent jump-diffusion models via integral equations and characteristic functions
Despite significant advancements in machine learning for derivative pricing, the efficient and accurate valuation of American options remains a persistent challenge due to complex exercise boundaries, near-expiry behavio…