paper-with-me

홈 › Papers

P3: Probabilistic Policy Propagation for Stable VAE-Based Robot Learning

2026-07-28 · Liyun Yan, Jianming Ma, Yang Zhang, Shengcheng Fu, Zhanxiang Cao, Keqi Zhu, Yizhi Chen, Yue Gao arxiv

Variational Autoencoders are widely used to encode high-dimensional and noisy observations in robotics. However, their stochastic latent creates a mismatch with Proximal Policy Optimization (PPO): an effective policy marginalizes over the latent distribution, whereas former implementations estimate its probability ratio and KL divergence using only one latent sample. We identify a fundamental but overlooked theoretical cause: naive single-sample approximations in stochastic latent space induce significant variance and bias in the surrogate loss. To address this, we introduce P^3 (Probabilistic Policy Propagation), a distribution-aware optimization framework for VAE-based policies. $P^3$ couples moment-based probabilistic method for stable and efficient learning with sampling-based calibration for robust policy behavior under latent uncertainty. In our experiments, P^3 boosts data efficiency from 64.6% to >96%, reduces convergence steps by >20%. Furthermore, P^3 is evaluated on challenging humanoid parkour tasks and shows an effective foundation for VAE-based PPO. Code is available at https://github.com/ylyem9x/P3_Open.

📄 PDF Abstract BibTeX arXiv:2607.25541

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Neuromorphic Reinforcement Learning for Quadruped Locomotion Control on Uneven Terrain

2026-05-10 · Zhuangyu Han, Abhronil Sengupta arxiv

Reinforcement learning (RL) has enabled robust quadruped locomotion over complex terrain, but most learned controllers are trained offline with backpropagation in massively parallel simulation and deployed as fixed polic…

Reinforcement Learning

Riemannian Motion Policy Fusion through Learnable Lyapunov Function Reshaping

2019-10-07 · Mustafa Mukadam, Ching-An Cheng, Dieter Fox, Byron Boots 외

RMPflow is a recently proposed policy-fusion framework based on differential geometry. While RMPflow has demonstrated promising performance, it requires the user to provide sensible subtask policies as Riemannian motion …

Imitation Learning

DiSProD: Differentiable Symbolic Propagation of Distributions for Planning

2023-02-03 · Palash Chatterjee, Ashutosh Chapagain, Weizhe Chen, Roni Khardon

The paper introduces DiSProD, an online planner developed for environments with probabilistic transitions in continuous state and action spaces. DiSProD builds a symbolic graph that captures the distribution of future tr…

Navigate

VINE: Taming Generative Control Policies for Reinforcement Learning

2026-07-11 · Rushuai Yang, Zhuo Han, Houlin Li, Hecheng Wang 외 arxiv

Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of complex and multimodal action distribut…

Reinforcement LearningOffline RL

Practical Probabilistic Model-based Deep Reinforcement Learning by Integrating Dropout Uncertainty and Trajectory Sampling

2023-09-20 · Wenjun Huang, Yunduan Cui, Huiyun Li, Xinyu Wu

This paper addresses the prediction stability, prediction accuracy and control capability of the current probabilistic model-based reinforcement learning (MBRL) built on neural networks. A novel approach dropout-based pr…

Deep Reinforcement LearningModel-based Reinforcement LearningMuJoCoPrediction