MDPs with Unawareness
Markov decision processes (MDPs) are widely used for modeling decision-making problems in robotics, automated control, and economics. Traditional MDPs assume that the decision maker (DM) knows all states and actions. However, this may not be true in many situations of interest. We define a new framework, MDPs with unawareness (MDPUs) to deal with the possibilities that a DM may not be aware of all possible actions. We provide a complete characterization of when a DM can learn to play near-optimally in an MDPU, and give an algorithm that learns to play near-optimally when it is possible to do so, as efficiently as possible. In particular, we characterize when a near-optimal solution can be found in polynomial time.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingSimilar Papers 제목 키워드 기반
MDPs with Unawareness in Robotics
We formalize decision-making problems in robotics and automated control using continuous MDPs and actions that take place over continuous time intervals. We then approximate the continuous MDP using finer and finer discr…
Decision MakingOn the state-space model of unawareness
We show that the knowledge of an agent carrying non-trivial unawareness violates the standard property of 'necessitation', therefore necessitation cannot be used to refute the standard state-space model. A revised versio…
modelKuhn's Theorem for Games of the Extensive Form with Unawareness
We extend Kuhn's Theorem to games of the extensive form with unawareness. This extension is not obvious: First, games of the extensive form with non-trivial unawareness involve a forest of partially ordered game trees ra…
FormRevisiting the state-space model of unawareness
We propose a knowledge operator based on the agent's possibility correspondence which preserves her non-trivial unawareness within the standard state-space model. Our approach may provide a solution to the classical impo…
modelRaising Bidders' Awareness in Second-Price Auctions
When bidders bid on complex objects, they might be unaware of characteristics effecting their valuations. We assume that each buyer's valuation is a sum of independent random variables, one for each characteristic. When …