paper-with-me

홈 › Papers

ISOPO: Proximal policy gradients without pi-old

2025-12-29 · Nilin Abrahamsen arxiv

This note introduces Isometric Policy Optimization (ISOPO), an efficient method to approximate the natural policy gradient in a single gradient step. In comparison, existing proximal policy methods such as GRPO or CISPO use multiple gradient steps with variants of importance ratio clipping to approximate a natural gradient step relative to a reference policy. In its simplest form, ISOPO normalizes the log-probability gradient of each sequence in the Fisher metric before contracting with the advantages. Another variant of ISOPO transforms the microbatch advantages based on the neural tangent kernel in each layer. ISOPO applies this transformation layer-wise in a single backward pass and can be implemented with negligible computational overhead compared to vanilla REINFORCE.

📄 PDF Abstract BibTeX arXiv:2512.23353

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gradient Informed Proximal Policy Optimization

2023-12-14 · NeurIPS 2023 11 · Sanghyun Son, Laura Yu Zheng, Ryan Sullivan, Yi-Ling Qiao 외

We introduce a novel policy learning method that integrates analytical gradients from differentiable environments with the Proximal Policy Optimization (PPO) algorithm. To incorporate analytical gradients into the PPO fr…

Natural Policy Gradients In Reinforcement Learning Explained

2022-09-05 · W. J. A. van Heeswijk

Traditional policy gradient methods are fundamentally flawed. Natural gradients converge quicker and better, forming the foundation of contemporary Reinforcement Learning such as Trust Region Policy Optimization (TRPO) a…

Policy Gradient Methodsreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Taking gradients through experiments: LSTMs and memory proximal policy optimization for black-box quantum control

2018-02-12 · Moritz August, José Miguel Hernández-Lobato

In this work we introduce the application of black-box quantum control as an interesting rein- forcement learning problem to the machine learning community. We analyze the structure of the reinforcement learning problems…

Reinforcement Learning

Dynamical System Optimization

2025-06-10 · Emo Todorov

We develop an optimization framework centered around a core idea: once a (parametric) policy is specified, control authority is transferred to the policy, resulting in an autonomous dynamical system. Thus we should be ab…

reinforcement-learningReinforcement Learning

Model-free Policy Learning with Reward Gradients

2021-03-09 · Qingfeng Lan, Samuele Tosatto, Homayoon Farrahi, A. Rupam Mahmood

Despite the increasing popularity of policy gradient methods, they are yet to be widely utilized in sample-scarce applications, such as robotics. The sample efficiency could be improved by making best usage of available …

Continuous ControlmodelMuJoCoPolicy Gradient Methods