paper-with-me

홈 › Papers

TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback

2025-08-04 · Lei Pang, Jun Luo, Ruinan Jin arxiv

Group Relative Policy Optimization (GRPO), recently introduced by DeepSeek, is a critic-free reinforcement learning algorithm for fine-tuning large language models. GRPO replaces the value function in Proximal Policy Optimization (PPO) with group-normalized rewards while retaining PPO-style token-level importance sampling based on an old policy. Our theoretical analysis reveals that the GRPO update rule estimates the policy gradient at the old policy rather than the current one; however, since the old policy is refreshed every few steps, the resulting discrepancy remains small and the induced bias is negligible in practice. To empirically validate this insight, we conduct an ablation study that entirely removes importance sampling and performs multiple optimization steps using gradients estimated at a fixed old policy. Remarkably, this simplified variant attains performance comparable to standard GRPO. Motivated by this finding, we propose Trajectory-level Importance-Corrected GRPO (TIC-GRPO), a new algorithm that replaces token-level importance ratios with a single trajectory-level probability ratio, thereby yielding an estimate of the current policy gradient while preserving the critic-free structure. Furthermore, we present the first convergence analysis for GRPO-style methods and show that TIC-GRPO converges faster than GRPO. Finally, empirical results across math reasoning and coding tasks demonstrate the superiority of TIC-GRPO.

📄 PDF Abstract BibTeX arXiv:2508.02833

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning

2025-10-16 · Xiangyu Meng, Zixian Zhang, Zhenghao Zhang, Junchao Liao 외 arxiv

While advanced methods like VACE and Phantom have advanced video generation for specific subjects in diverse scenarios, they struggle with multi-human identity preservation in dynamic interactions, where consistent ident…

Reinforcement LearningVideo Generation

DanceGRPO: Unleashing GRPO on Visual Generation

2025-05-12 · Zeyue Xue, Jie Wu, Yu Gao, Fangyuan Kong 외

Recent breakthroughs in generative models-particularly diffusion models and rectified flows-have revolutionized visual content creation, yet aligning model outputs with human preferences remains a critical challenge. Exi…

Denoisingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

2025-11-11 · Chanakya Ekbote, Vijay Lingam, Sujay Sanghavi, Jun Huan 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard recipe for post-training LLMs on reasoning tasks, with Group Relative Policy Optimization (GRPO) emerging as a leading approach. However, GRPO a…

Reinforcement LearningCode Generation

Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment

2026-02-17 · Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi 외 arxiv

Despite their sophisticated general-purpose capabilities, Large Language Models (LLMs) often fail to align with diverse individual preferences because standard post-training methods, like Reinforcement Learning with Huma…

Reinforcement Learning

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

2025-05-29 · Mingzhe Du, Luu Anh Tuan, Yue Liu, Yuhao QING 외

Large Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we introduce a novel test-time iterative optim…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)