paper-with-me

홈 › Papers

KLip-PPO: A per-sample KL perspective on PPO-Clip

2026-06-22 · Riccardo Colletti, Robin Holzinger arxiv

Proximal Policy Optimization (PPO) is the standard policy-gradient algorithm for on-policy reinforcement learning. The literature presents it in two forms, a clipped surrogate that bounds the importance ratio between successive policies and a Kullback-Leibler penalty between them. These forms are treated as separate algorithms with their own gradients, their own hyperparameters, and their own reference implementations, and a sizeable body of empirical work compares them. We show that the gradient of the clipped surrogate is reproduced exactly by a Kullback-Leibler surrogate whose coefficient varies per sample, with closed-form dependence on the importance ratio and the advantage. The identity holds at every minibatch step and across the entire inner loop, and on five MuJoCo continuous-control benchmarks the two losses produce indistinguishable training curves. The reformulation exposes a structural feature of the clipped surrogate that the min notation hides. PPO-Clip's implicit per-sample penalty is a step function at the boundary of the trust region, and the shape of this coefficient is the natural design axis for generalising the algorithm. We sketch the resulting follow-up directions in the discussion.

📄 PDF Abstract BibTeX arXiv:2606.23932

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation

2025-05-08 · Haokun Lin, Teng Wang, Yixiao Ge, Yuying Ge 외

Pioneering token-based works such as Chameleon and Emu3 have established a foundation for multimodal unification but face challenges of high training computational overhead and limited comprehension performance due to a …

Quantization

SPKLIP: Aligning Spike Video Streams with Natural Language

2025-05-19 · Yongchang Gao, Meiling Jin, Zhaofei Yu, Tiejun Huang 외

Spike cameras offer unique sensing capabilities but their sparse, asynchronous output challenges semantic understanding, especially for Spike Video-Language Alignment (Spike-VLA) where models like CLIP underperform due t…

Contrastive LearningFew-Shot Learning

RetailKLIP : Finetuning OpenCLIP backbone using metric learning on a single GPU for Zero-shot retail product image classification

2023-12-16 · Muktabh Mayank Srivastava

Retail product or packaged grocery goods images need to classified in various computer vision applications like self checkout stores, supply chain automation and retail execution evaluation. Previous works explore ways t…

GPUimage-classificationImage ClassificationMetric Learning

KLIP: localized distribution shift detection via KL-divergence with diffusion priors in Inverse Problems

2026-05-29 · Alireza Kheirandish, Jihoon Hong, Sara Fridovich-Keil arxiv

Diffusion models have shown promising performance as data-driven priors for computational imaging, as well as some capacity to detect out-of-distribution (OOD) images. However, existing approaches to OOD detection often …

Karhunen-Loève Data Imputation in High Contrast Imaging

2023-08-31 · Bin B. Ren

Detection and characterization of extended structures is a crucial goal in high contrast imaging. However, these structures face challenges in data reduction, leading to over-subtraction from speckles and self-subtractio…

Imputation