paper-with-me

Papers

Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments

2026-03-06 · Ege C. Kaya, Mahsa Ghasemi, Abolfazl Hashemi arxiv

Many distributional quantities in reinforcement learning are intrinsically joint across actions, including distributions of gaps and probabilities of superiority. However, the classical Markov decision process (MDP) formalism specifies only marginal laws and leaves the joint law of counterfactual one-step outcomes across multiple possible actions at a state unspecified. We study coupled-dynamics environments with a multi-action generative interface which can sample counterfactual one-step outcomes for multiple actions under shared exogenous randomness. We propose joint MDPs (JMDPs) as a formalism for such environments by augmenting an MDP with a multi-action sample transition model which specifies a coupling of one-step counterfactual outcomes, while preserving standard MDP interaction as marginal observations. We adopt and formalize a one-step coupling regime where dependence across actions is confined to immediate counterfactual outcomes at the queried state. In this regime, we derive Bellman operators for $n$th-order return moments, providing dynamic programming and incremental algorithms with convergence guarantees.

📄 PDF Abstract BibTeX arXiv:2603.06946

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes

2026-08-24 · Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi, Vijay Gupta 외 arxiv

Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The ordinary Markov decision process formalism…

Off-Policy Evaluation in Partially Observable Environments

2019-09-09 · Guy Tennenholtz, Shie Mannor, Uri Shalit

This work studies the problem of batch off-policy evaluation for Reinforcement Learning in partially observable environments. Off-policy evaluation under partial observability is inherently prone to bias, with risk of ar…

Off-policy evaluationReinforcement LearningReinforcement Learning (RL)

Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling

2025-08-06 · Yongyi Wang, Lingfeng Li, Bozhou Chen, Ang Li 외 arxiv

Recent benchmarks for memory-augmented reinforcement learning (RL) have introduced partially observable Markov decision process (POMDP) environments in which agents must use historical observations to make decisions. How…

Reinforcement Learning

Reinforcement Learning with History-Dependent Dynamic Contexts

2023-02-04 · Guy Tennenholtz, Nadav Merlis, Lior Shani, Martin Mladenov 외

We introduce Dynamic Contextual Markov Decision Processes (DCMDPs), a novel reinforcement learning framework for history-dependent environments that generalizes the contextual MDP framework to handle non-Markov environme…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Budgeted Reinforcement Learning in Continuous State Space

2019-03-03 · NeurIPS 2019 12 · Nicolas Carrara, Edouard Leurent, Romain Laroche, Tanguy Urvoy 외

A Budgeted Markov Decision Process (BMDP) is an extension of a Markov Decision Process to critical applications requiring safety constraints. It relies on a notion of risk implemented in the shape of a cost signal constr…

Autonomous DrivingDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1