Monte Carlo Information-Oriented Planning
In this article, we discuss how to solve information-gathering problems expressed as rho-POMDPs, an extension of Partially Observable Markov Decision Processes (POMDPs) whose reward rho depends on the belief state. Point-based approaches used for solving POMDPs have been extended to solving rho-POMDPs as belief MDPs when its reward rho is convex in B or when it is Lipschitz-continuous. In the present paper, we build on the POMCP algorithm to propose a Monte Carlo Tree Search for rho-POMDPs, aiming for an efficient on-line planner which can be used for any rho function. Adaptations are required due to the belief-dependent rewards to (i) propagate more than one state at a time, and (ii) prevent biases in value estimates. An asymptotic convergence proof to epsilon-optimal values is given when rho is continuous. Experiments are conducted to analyze the algorithms at hand and show that they outperform myopic approaches.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Prompt-Based Monte-Carlo Tree Search for Goal-Oriented Dialogue Policy Planning
Planning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress. Many approaches thus consider training neural networks to perform look-ahead search algorithms such…
Language ModelingLanguage ModellingLarge Language ModelFaithful Question Answering with Monte-Carlo Planning
Although large language models demonstrate remarkable question-answering performances, revealing the intermediate reasoning steps that the models faithfully follow remains challenging. In this paper, we propose FAME (FAi…
Decision MakingQuestion AnsweringvalidFeedback-Aware Monte Carlo Tree Search for Efficient Information Seeking in Goal-Oriented Conversations
Effective decision-making and problem-solving in conversational systems require the ability to identify and acquire missing information through targeted questioning. A key challenge lies in efficiently narrowing down a l…
Medical DiagnosisSemantic SimilaritySemantic Textual SimilarityProbabilistic Planning with Sequential Monte Carlo methods
In this work, we propose a novel formulation of planning which views it as a probabilistic inference problem over future optimal trajectories. This enables us to use sampling methods, and thus, tackle planning in continu…
continuous-controlContinuous ControlSensitivity Analyses of Resilience-oriented Risk-averse Active Distribution Systems Planning
This paper presents sensitivity analyses of resilience-based active distribution system planning solutions with respect to different parameters. The distribution system planning problem is formulated as a two-stage risk-…
SensitivityStochastic Optimization