paper-with-me

Papers

Stick-Breaking Policy Learning in Dec-POMDPs

2015-05-01 · Miao Liu, Christopher Amato, Xuejun Liao, Lawrence Carin, Jonathan P. How

Expectation maximization (EM) has recently been shown to be an efficient algorithm for learning finite-state controllers (FSCs) in large decentralized POMDPs (Dec-POMDPs). However, current methods use fixed-size FSCs and often converge to maxima that are far from optimal. This paper considers a variable-size FSC to represent the local policy of each agent. These variable-size FSCs are constructed using a stick-breaking prior, leading to a new framework called \emph{decentralized stick-breaking policy representation} (Dec-SBPR). This approach learns the controller parameters with a variational Bayesian algorithm without having to assume that the Dec-POMDP model is available. The performance of Dec-SBPR is demonstrated on several benchmark problems, showing that the algorithm scales to large problems while outperforming other state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:1505.00274

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Robustness Analysis of POMDP Policies to Observation Perturbations

2026-04-23 · Benjamin Kraske, Qi Heng Ho, Federico Rossi, Morteza Lahijanian 외 arxiv

Policies for Partially Observable Markov Decision Processes (POMDPs) are often designed using a nominal system model. In practice, this model can deviate from the true system during deployment due to factors such as cali…

An elementary derivation of the Chinese restaurant process from Sethuraman's stick-breaking process

2018-01-01 · Jeffrey W. Miller

The Chinese restaurant process (CRP) and the stick-breaking process are the two most commonly used representations of the Dirichlet process. However, the usual proof of the connection between them is indirect, relying on…

Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study

2024-10-23 · Shawn Tan, Songlin Yang, Aaron Courville, Rameswar Panda 외

The self-attention mechanism traditionally relies on the softmax operator, necessitating positional embeddings like RoPE, or position biases to account for token order. But current methods using still face length general…

Stick-Breaking Variational Autoencoders

2016-05-20 · Eric Nalisnick, Padhraic Smyth

We extend Stochastic Gradient Variational Bayes to perform posterior inference for the weights of Stick-Breaking processes. This development allows us to define a Stick-Breaking Variational Autoencoder (SB-VAE), a Bayesi…

Tree-Structured Stick Breaking for Hierarchical Data

2010-12-01 · NeurIPS 2010 12 · Zoubin Ghahramani, Michael. I. Jordan, Ryan P. Adams

Many data are naturally modeled by an unobserved hierarchical structure. In this paper we propose a flexible nonparametric prior over unknown data hierarchies. The approach uses nested stick-breaking processes to allow f…

Bayesian InferenceClustering