paper-with-me

Papers

Scalar-Stepsize Nonuniform Monte Carlo Optimistic Policy Iteration: A Certified Counterexample

2026-06-14 · Yuanlong Chen arxiv

Tsitsiklis proved convergence of Monte Carlo optimistic policy iteration under a uniform update structure and identified nonuniform update frequencies as a delicate obstruction. We give a certified negative answer for the natural scalar-stepsize, unnormalized asynchronous state-value recursion with fixed nonuniform state-selection probabilities. In a three-state, two-action discounted MDP, the nonuniform update frequencies induce a diagonally scaled greedy-policy mean field with a certified nonconstant attracting hybrid periodic orbit. With a bounded unbiased geometric-horizon estimator and Robbins--Monro stepsizes, the original stochastic recursion remains trapped near the cycle with positive probability and therefore fails to converge. The example pinpoints a geometric obstruction: uniform sampling gives radial residual contraction, whereas scalar nonuniform sampling anisotropically distorts the residual dynamics and can generate switched attracting cycles.

📄 PDF Abstract BibTeX arXiv:2606.15978

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Regime-Switching Langevin Monte Carlo Algorithms

2025-08-31 · Xiaoyu Wang, Yingli Wang, Lingjiong Zhu arxiv

Langevin Monte Carlo (LMC) algorithms are popular Markov Chain Monte Carlo (MCMC) methods to sample a target probability distribution, which arises in many applications in machine learning. Inspired by regime-switching s…

On the convergence of optimistic policy iteration for stochastic shortest path problem

2018-08-27 · Yuanlong Chen

In this paper, we prove some convergence results of a special case of optimistic policy iteration algorithm for stochastic shortest path problem. We consider both Monte Carlo and $TD(\lambda)$ methods for the policy eval…

Decentralized Bayesian Learning with Metropolis-Adjusted Hamiltonian Monte Carlo

2021-07-15 · Vyacheslav Kungurtsev, Adam Cobb, Tara Javidi, Brian Jalaian

Federated learning performed by a decentralized networks of agents is becoming increasingly important with the prevalence of embedded software on autonomous devices. Bayesian approaches to learning benefit from offering …

Federated Learning

Limited depth bandit-based strategy for Monte Carlo planning in continuous action spaces

2021-06-29 · Ricardo Quinteiro, Francisco S. Melo, Pedro A. Santos

This paper addresses the problem of optimal control using search trees. We start by considering multi-armed bandit problems with continuous action spaces and propose LD-HOO, a limited depth variant of the hierarchical op…

Convergence of Monte Carlo Optimistic Policy Iteration: Beyond Uniform State-Action Updates

2026-06-09 · Octave Oliviers, Glenn Vinnicombe arxiv

The asymptotic behaviour of Monte Carlo optimistic policy iteration (MC-O-PI) is a long-standing open question. When the model of the environment is unknown, as is common in practice, the only known condition that guaran…