paper-with-me

홈 › Papers

Constrained Group Relative Policy Optimization

2026-02-05 · Roger Girgis, Rodrigue de Schaetzen, Luke Rowe, Azalée Robitaille, Christopher Pal, Liam Paull arxiv

Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with constrained policy optimization (e.g. for safety-critical domains) has not been carefully examined. In this work, we introduce Constrained GRPO, a Lagrangian-based extension of GRPO for constrained policy optimization. We show that the standard practice of scalarizing rewards before normalization introduces a critical Lagrangian-specific failure mode: GRPO's within-group normalization makes constrained optimization highly sensitive to how multi-component learning signals are aggregated. We show that scalarizing rewards before normalization introduces shared-denominator coupling, so that changing one multiplier alters not only the emphasis on its corresponding constraint, but also the relative weighting of the reward and other constraints. We address this with a simple but crucial modification: scalarizing standardized advantages rather than rewards. This yields a better-conditioned update by addressing the coupling induced by reward scalarization, resulting in better-behaved multiplier dynamics and more stable constraint enforcement in practice. Empirically, across a controlled gridworld, a real-world autonomous driving benchmark, and a mathematical reasoning task, Constrained GRPO consistently achieves better adherence to specified constraints while maintaining or improving task performance.

📄 PDF Abstract BibTeX arXiv:2602.05863

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GVPO: Group Variance Policy Optimization for Large Language Model Post-Training

2025-04-28 · Kaichen Zhang, Yuzhong Hong, Junwei Bao, Hongfei Jiang 외

Post-training plays a crucial role in refining and aligning large language models to meet specific tasks and human preferences. While recent advancements in post-training techniques, such as Group Relative Policy Optimiz…

Language ModelingLanguage ModellingLarge Language Model

MC-GRPO: Median-Centered Group Relative Policy Optimization for Small-Rollout Reinforcement Learning

2026-01-30 · Youngeun Kim arxiv

Group-relative policy optimization methods train language models by generating multiple rollouts per prompt and normalizing rewards with a shared mean reward baseline. In resource-constrained settings where the rollout b…

Reinforcement Learning

Advancing SLM Tool-Use Capability using Reinforcement Learning

2025-09-03 · Dhruvi Paprunia, Vansh Kharidia, Pankti Doshi arxiv

In an era where tool-augmented AI agents are becoming increasingly vital, our findings highlight the ability of Group Relative Policy Optimization (GRPO) to empower SLMs, which are traditionally constrained in tool use. …

Reinforcement Learning

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models

2026-04-18 · Zhaokang Liao, Yingguo Gao, Yi Yang, Yongheng Hu 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising approach to improve the reasoning abilities of Large Language Models (LLMs). Among RLVR algorithms, Group Relative Policy Optimization (GRP…

Reinforcement Learning

Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing

2026-06-01 · Azal Ahmad Khan, Ammar Ahmed, Zeshan Fayyaz, Sheng Di 외 arxiv

Synchronous reinforcement learning methods such as Group Relative Policy Optimization (GRPO) provide stable and reproducible on-policy training, but they are highly vulnerable to stragglers, a single unusually long rollo…

Reinforcement Learning