paper-with-me

Papers

Verifiable RNN-Based Policies for POMDPs Under Temporal Logic Constraints

2020-02-13 · Steven Carr, Nils Jansen, Ufuk Topcu

Recurrent neural networks (RNNs) have emerged as an effective representation of control policies in sequential decision-making problems. However, a major drawback in the application of RNN-based policies is the difficulty in providing formal guarantees on the satisfaction of behavioral specifications, e.g. safety and/or reachability. By integrating techniques from formal methods and machine learning, we propose an approach to automatically extract a finite-state controller (FSC) from an RNN, which, when composed with a finite-state system model, is amenable to existing formal verification tools. Specifically, we introduce an iterative modification to the so-called quantized bottleneck insertion technique to create an FSC as a randomized policy with memory. For the cases in which the resulting FSC fails to satisfy the specification, verification generates diagnostic information. We utilize this information to either adjust the amount of memory in the extracted FSC or perform focused retraining of the RNN. While generally applicable, we detail the resulting iterative procedure in the context of policy synthesis for partially observable Markov decision processes (POMDPs), which is known to be notoriously hard. The numerical experiments show that the proposed approach outperforms traditional POMDP synthesis methods by 3 orders of magnitude within 2% of optimal benchmark values.

📄 PDF Abstract BibTeX arXiv:2002.05615

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingDiagnosticSequential Decision Making

Similar Papers 제목 키워드 기반

Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial Observability

2025-10-27 · Eline M. Bovy, Caleb Probine, Marnix Suilen, Ufuk Topcu 외 arxiv

Multi-environment POMDPs (ME-POMDPs) extend standard POMDPs with discrete model uncertainty. ME-POMDPs represent a finite set of POMDPs that share the same state, action, and observation spaces, but may arbitrarily vary …

\textsc{rfPG}: Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs

2025-05-14 · Maris F. L. Galesloot, Roman Andriushchenko, Milan Češka, Sebastian Junges 외

Partially observable Markov decision processes (POMDPs) model specific environments in sequential decision-making under uncertainty. Critically, optimal policies for POMDPs may not be robust against perturbations in the …

Decision Making Under UncertaintySequential Decision Making

Finite-State Controllers for (Hidden-Model) POMDPs using Deep Reinforcement Learning

2026-02-09 · David Hudák, Maris F. L. Galesloot, Martin Tappler, Martin Kurečka 외 arxiv

Solving partially observable Markov decision processes (POMDPs) requires computing policies under imperfect state information. Despite recent advances, the scalability of existing POMDP solvers remains limited. Moreover,…

Reinforcement Learning

Evidential Transactions with Cyberlogic

2023-03-20 · Harald Ruess, Natarajan Shankar

Cyberlogic is an enabling logical foundation for building and analyzing digital transactions that involve the exchange of digital forms of evidence. It is based on an extension of (first-order) intuitionistic predicate l…

Agent policies from higher-order causal functions

2025-12-11 · Matt Wilson arxiv

We establish a correspondence between equivalence classes of agent-state policies for deterministic POMDPs and one-input process functions (the classical-deterministic limit of higher-order quantum operations). We use th…