paper-with-me

Papers

Value Functions for Temporal Logic: Optimal Policies and Safety Filters

2026-05-01 · Oswin So, William Sharpless, Sylvia Herbert, Chuchu Fan arxiv

While Bellman equations for basic reach, avoid, and reach-avoid problems are well studied, the relationship between value optimality and policy optimality becomes subtle in the undiscounted infinite-horizon setting, particularly for more complicated tasks. Greedily maximizing the Q-function can produce policies that indefinitely defer task completion for reach-avoid problems, or equivalently, Until specifications, even when the value function is optimal. Building upon recent results decomposing the value function for temporal logic (TL) into a graph of constituent value functions, we construct non-Markovian policies based on state history that avoid this pathology and prove their optimality with respect to the quantitative robustness score for nested Until, Globally, and Globally-Until specifications. We further show how the Q function can serve as a safety filter for complex TL specifications, extending prior results beyond simple avoid or reach-avoid tasks.

📄 PDF Abstract BibTeX arXiv:2605.01051

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Constraint Satisfaction Propagation: Non-stationary Policy Synthesis for Temporal Logic Planning

2019-01-29 · Thomas J. Ringstrom, Paul R. Schrater

Problems arise when using reward functions to capture dependencies between sequential time-constrained goal states because the state-space must be prohibitively expanded to accommodate a history of successfully achieved …

Guaranteed Completion of Complex Tasks via Temporal Logic Trees and Hamilton-Jacobi Reachability

2024-04-12 · Frank J. Jiang, Kaj Munhoz Arfvidsson, Chong He, Mo Chen 외

In this paper, we present an approach for guaranteeing the completion of complex tasks with cyber-physical systems (CPS). Specifically, we leverage temporal logic trees constructed using Hamilton-Jacobi reachability anal…

Learning General Optimal Policies with Graph Neural Networks: Expressive Power, Transparency, and Limits

2021-09-21 · Simon Ståhlberg, Blai Bonet, Hector Geffner

It has been recently shown that general policies for many classical planning domains can be expressed and learned in terms of a pool of features defined from the domain predicates using a description logic grammar. At th…

Combinatorial Optimization

Compositional planning in Markov decision processes: Temporal abstraction meets generalized logic composition

2018-10-05 · Xuan Liu, Jie Fu

In hierarchical planning for Markov decision processes (MDPs), temporal abstraction allows planning with macro-actions that take place at different time scale in form of sequential composition. In this paper, we propose …

Model-Agnostic Solutions for Deep Reinforcement Learning in Non-Ergodic Contexts

2026-01-13 · Bert Verbruggen, Arne Vanhoyweghen, Vincent Ginis arxiv

Reinforcement Learning (RL) remains a central optimisation framework in machine learning. Although RL agents can converge to optimal solutions, the definition of ``optimality'' depends on the environment's statistical pr…

Reinforcement Learning