paper-with-me

Papers

Equilibrium in Misspecified Markov Decision Processes

2016-05-14

We study Markov decision problems where the agent does not know the transition probability function mapping current states and actions to future states. The agent has a prior belief over a set of possible transition functions and updates beliefs using Bayes' rule. We allow her to be misspecified in the sense that the true transition probability function is not in the support of her prior. This problem is relevant in many economic settings but is usually not amenable to analysis by the researcher. We make the problem tractable by studying asymptotic behavior. We propose an equilibrium notion and provide conditions under which it characterizes steady state behavior. In the special case where the problem is static, equilibrium coincides with the single-agent version of Berk-Nash equilibrium (Esponda and Pouzo (2016)). We also discuss subtle issues that arise exclusively in dynamic settings due to the possibility of a negative value of experimentation.

📄 PDF Abstract BibTeX arXiv:1502.06901

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On Existence of Berk-Nash Equilibria in Misspecified Markov Decision Processes with Infinite Spaces

2022-06-16 · Robert M. Anderson, Haosui Duanmu, Aniruddha Ghosh, M. Ali Khan

Model misspecification is a critical issue in many areas of theoretical and empirical economics. In the specific context of misspecified Markov Decision Processes, Esponda and Pouzo (2021) defined the notion of Berk-Nash…

Iterative Hierarchical Optimization for Misspecified Problems (IHOMP)

2016-02-10 · Daniel J. Mankowitz, Timothy A. Mann, Shie Mannor

For complex, high-dimensional Markov Decision Processes (MDPs), it may be necessary to represent the policy with function approximation. A problem is misspecified whenever, the representation cannot express any policy wi…

Robust $Q$-learning Algorithm for Markov Decision Processes under Wasserstein Uncertainty

2022-09-30 · Ariel Neufeld, Julian Sester

We present a novel $Q$-learning algorithm tailored to solve distributionally robust Markov decision problems where the corresponding ambiguity set of transition probabilities for the underlying Markov decision process is…

Q-Learning

Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs

2023-10-13 · Debangshu Banerjee, Aditya Gopalan

Parametric, feature-based reward models are employed by a variety of algorithms in decision-making settings such as bandits and Markov decision processes (MDPs). The typical assumption under which the algorithms are anal…

Decision MakingMulti-Armed BanditsQ-Learning

Approximate State Abstraction for Markov Games

2024-12-20 · Hiroki Ishibashi, Kenshi Abe, Atsushi Iwasaki

This paper introduces state abstraction for two-player zero-sum Markov games (TZMGs), where the payoffs for the two players are determined by the state representing the environment and their respective actions, with stat…