paper-with-me

홈 › Papers

MDPs with Unawareness in Robotics

2020-05-20 · Nan Rong, Joseph Y. Halpern, Ashutosh Saxena

We formalize decision-making problems in robotics and automated control using continuous MDPs and actions that take place over continuous time intervals. We then approximate the continuous MDP using finer and finer discretizations. Doing this results in a family of systems, each of which has an extremely large action space, although only a few actions are "interesting". We can view the decision maker as being unaware of which actions are "interesting". We can model this using MDPUs, MDPs with unawareness, where the action space is much smaller. As we show, MDPUs can be used as a general framework for learning tasks in robotic problems. We prove results on the difficulty of learning a near-optimal policy in an an MDPU for a continuous task. We apply these ideas to the problem of having a humanoid robot learn on its own how to walk.

📄 PDF Abstract BibTeX arXiv:2005.10381

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

MDPs with Unawareness

2014-07-27 · Joseph Y. Halpern, Nan Rong, Ashutosh Saxena

Markov decision processes (MDPs) are widely used for modeling decision-making problems in robotics, automated control, and economics. Traditional MDPs assume that the decision maker (DM) knows all states and actions. How…

Decision Making

On the state-space model of unawareness

2023-04-10 · Alex A. T. Rathke

We show that the knowledge of an agent carrying non-trivial unawareness violates the standard property of 'necessitation', therefore necessitation cannot be used to refute the standard state-space model. A revised versio…

model

Kuhn's Theorem for Games of the Extensive Form with Unawareness

2025-03-05 · Ki Vin Foo, Burkhard C. Schipper

We extend Kuhn's Theorem to games of the extensive form with unawareness. This extension is not obvious: First, games of the extensive form with non-trivial unawareness involve a forest of partially ordered game trees ra…

Form

Qualitative Possibilistic Mixed-Observable MDPs

2013-09-26 · Nicolas Drougard, Florent Teichteil-Konigsbuch, Jean-Loup Farges, Didier Dubois

Possibilistic and qualitative POMDPs (pi-POMDPs) are counterparts of POMDPs used to model situations where the agent's initial belief or observation probabilities are imprecise due to lack of past experiences or insuffic…

Memory-based Deep Reinforcement Learning for POMDPs

2021-02-24 · Lingheng Meng, Rob Gorbet, Dana Kulić

A promising characteristic of Deep Reinforcement Learning (DRL) is its capability to learn optimal policy in an end-to-end manner without relying on feature engineering. However, most approaches assume a fully observable…

Deep Reinforcement LearningFeature Engineeringreinforcement-learningReinforcement Learning+1