paper-with-me

Papers

Regular Decision Processes for Grid Worlds

2021-11-05 · Nicky Lenaers, Martijn van Otterlo

Markov decision processes are typically used for sequential decision making under uncertainty. For many aspects however, ranging from constrained or safe specifications to various kinds of temporal (non-Markovian) dependencies in task and reward structures, extensions are needed. To that end, in recent years interest has grown into combinations of reinforcement learning and temporal logic, that is, combinations of flexible behavior learning methods with robust verification and guarantees. In this paper we describe an experimental investigation of the recently introduced regular decision processes that support both non-Markovian reward functions as well as transition functions. In particular, we provide a tool chain for regular decision processes, algorithmic extensions relating to online, incremental learning, an empirical evaluation of model-free and model-based solution algorithms, and applications in regular, but non-Markovian, grid worlds.

📄 PDF Abstract BibTeX arXiv:2111.03647

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingDecision Making Under UncertaintyIncremental LearningSequential Decision Making

Similar Papers 제목 키워드 기반

A Blackbox Approach to Best of Both Worlds in Bandits and Beyond

2023-02-20 · Christoph Dann, Chen-Yu Wei, Julian Zimmert

Best-of-both-worlds algorithms for online learning which achieve near-optimal regret in both the adversarial and the stochastic regimes have received growing attention recently. Existing techniques often require careful …

Multi-Armed Bandits

Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes

2026-02-01 · Yu Chen, Yuhao Liu, Jiatai Huang, Yihan Du 외 arxiv

We investigate episodic Markov Decision Processes with heavy-tailed losses (HTMDPs). Existing approaches for HTMDPs are conservative in stochastic environments and lack adaptivity in adversarial regimes. In this work, we…

Flex-Convolution (Million-Scale Point-Cloud Learning Beyond Grid-Worlds)

2018-03-20 · Fabian Groh, Patrick Wieschollek, Hendrik P. A. Lensch

Traditional convolution layers are specifically designed to exploit the natural data representation of images -- a fixed and regular grid. However, unstructured data like 3D point clouds containing irregular neighborhood…

Classify 3D Point CloudsGPUSemantic Segmentation

Best of Both Worlds Policy Optimization

2023-02-18 · Christoph Dann, Chen-Yu Wei, Julian Zimmert

Policy optimization methods are popular reinforcement learning algorithms in practice. Recent works have built theoretical foundation for them by proving $\sqrt{T}$ regret bounds even when the losses are adversarial. Suc…

Simultaneously Learning Stochastic and Adversarial Episodic MDPs with Known Transition

2020-06-10 · NeurIPS 2020 12 · Tiancheng Jin, Haipeng Luo

This work studies the problem of learning episodic Markov Decision Processes with known transition and bandit feedback. We develop the first algorithm with a ``best-of-both-worlds'' guarantee: it achieves $\mathcal{O}(lo…

Multi-Armed Bandits