paper-with-me

홈 › Papers

On Wasserstein Reinforcement Learning and the Fokker-Planck equation

2017-12-19 · Pierre H. Richemond, Brendan Maginnis

Policy gradients methods often achieve better performance when the change in policy is limited to a small Kullback-Leibler divergence. We derive policy gradients where the change in policy is limited to a small Wasserstein distance (or trust region). This is done in the discrete and continuous multi-armed bandit settings with entropy regularisation. We show that in the small steps limit with respect to the Wasserstein distance $W_2$, policy dynamics are governed by the Fokker-Planck (heat) equation, following the Jordan-Kinderlehrer-Otto result. This means that policies undergo diffusion and advection, concentrating near actions with high reward. This helps elucidate the nature of convergence in the probability matching setup, and provides justification for empirical practices such as Gaussian policy priors and additive gradient noise.

📄 PDF Abstract BibTeX arXiv:1712.07185

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Closing the ODE-SDE gap in score-based diffusion models through the Fokker-Planck equation

2023-11-27 · Teo Deveney, Jan Stanczuk, Lisa Maria Kreusser, Chris Budd 외

Score-based diffusion models have emerged as one of the most promising frameworks for deep generative modelling, due to their state-of-the art performance in many generation tasks while relying on mathematical foundation…

Large-Scale Wasserstein Gradient Flows

2021-06-01 · NeurIPS 2021 12 · Petr Mokrov, Alexander Korotin, Lingxiao Li, Aude Genevay 외

Wasserstein gradient flows provide a powerful means of understanding and solving many diffusion equations. Specifically, Fokker-Planck equations, which model the diffusion of probability measures, can be understood as gr…

Deterministic Fokker-Planck Transport -- With Applications to Sampling, Variational Inference, Kernel Mean Embeddings & Sequential Monte Carlo

2024-10-11 · Ilja Klebanov

The Fokker-Planck equation can be reformulated as a continuity equation, which naturally suggests using the associated velocity field in particle flow methods. While the resulting probability flow ODE offers appealing pr…

Density EstimationVariational Inference

Self-Consistency of the Fokker-Planck Equation

2022-06-02 · Zebang Shen, Zhenfu Wang, Satyen Kale, Alejandro Ribeiro 외

The Fokker-Planck equation (FPE) is the partial differential equation that governs the density evolution of the It\^o process and is of great importance to the literature of statistical physics and machine learning. The …

Score-based Transport Modeling for Mean-Field Fokker-Planck Equations

2023-04-21 · Jianfeng Lu, Yue Wu, Yang Xiang

We use the score-based transport modeling method to solve the mean-field Fokker-Planck equations, which we call MSBTM. We establish an upper bound on the time derivative of the Kullback-Leibler (KL) divergence to MSBTM n…