Optimally Solving Two-Agent Decentralized POMDPs Under One-Sided Information Sharing
Optimally solving decentralized partially observable Markov decision processes under either full or no information sharing received significant attention in recent years. However, little is known about how partial information sharing affects existing theory and algorithms. This paper addresses this question for a team of two agents, with one-sided information sharing---\ie both agents have imperfect information about the state of the world, but only one has access to what the other sees and does. From the perspective of a central planner, we show that the original problem can be reformulated into an equivalent information-state Markov decision process and solved as such. Besides, we prove that the optimal value function exhibits a specific form of uniform continuity. We also present a heuristic search algorithm utilizing this property and providing the first results for this family of problems.
Code (0)
등록된 구현이 없습니다.
Tasks
Heuristic SearchVocal Bursts Valence PredictionSimilar Papers 제목 키워드 기반
Optimally Solving Two-Agent Decentralized POMDPs Under One-Sided Information Sharing
Optimally solving decentralized partially observable Markov decision processes under either full or no information sharing received significant attention in recent years. However, little is known about how partial inform…
Heuristic SearchVocal Bursts Valence PredictionOptimally Solving Simultaneous-Move Dec-POMDPs: The Sequential Central Planning Approach
The centralized training for decentralized execution paradigm emerged as the state-of-the-art approach to $\epsilon$-optimally solving decentralized partially observable Markov decision processes. However, scalability re…
Information Gathering in Decentralized POMDPs by Policy Graph Improvement
Decentralized policies for information gathering are required when multiple autonomous agents are deployed to collect data about a phenomenon of interest without the ability to communicate. Decentralized partially observ…
Decision MakingQualitative Possibilistic Mixed-Observable MDPs
Possibilistic and qualitative POMDPs (pi-POMDPs) are counterparts of POMDPs used to model situations where the agent's initial belief or observation probabilities are imprecise due to lack of past experiences or insuffic…
Incremental Clustering and Expansion for Faster Optimal Planning in Dec-POMDPs
This article presents the state-of-the-art in optimal solution methods for decentralized partially observable Markov decision processes (Dec-POMDPs), which are general models for collaborative multiagent planning under u…
Clustering