paper-with-me

홈 › Papers

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning

2026-07-05 · Mingxuan Fan, Peiyang Liu arxiv

Group-based reinforcement learning (RL) has become an effective paradigm for improving large language model agents on long-horizon interactive tasks. To obtain finer-grained policy updates than trajectory-level optimization, recent work has moved toward step-level group-based RL, where intermediate steps are grouped and compared within a rollout batch. However, step-level advantage estimation is sensitive to how groups are formed: grouping by broad state keys improves coverage but may compare actions taken under different histories, while enforcing historical consistency yields fairer comparisons at the cost of fragmented groups and missing peer-comparison signal. In this paper, we propose ProGPO (Progress- and Reliability-Oriented Group Policy Optimization), a learned-critic-free method for context-consistent step-level learning. ProGPO keeps exact-prefix action comparison, and complements sparse peer comparisons with transition credit derived from rollout-based state potentials. To estimate these potentials reliably, ProGPO combines semantic expansion with inverse-variance fusion across history depths. We evaluate ProGPO on two challenging agentic tasks, ALFWorld and WebShop, with Qwen2.5-1.5B-Instruct. Results show that ProGPO improves over matched agentic RL baselines under comparable computational overhead, and additional Qwen2.5-3B-Instruct experiments further test the scalability of the proposed method.

📄 PDF Abstract BibTeX arXiv:2607.04242

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents

2026-08-20 · Bo Qian, Yuting Wu, Shuang Zeng, Huaiyu Wan 외 arxiv

Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level signals into step-level credits through ste…

Reinforcement Learning

Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing

2026-04-02 · Gengsheng Li, Tianyu Yang, Junfeng Fang, Mingyang Song 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models. While Group Relative Policy Optimization (GRPO) is widely adopted, its coarse credit assignmen…

Reinforcement Learning

GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models

2026-08-18 · Peizheng Guo, Jianqi Zhang, Xingyu Zhang, Yun Fan 외 arxiv

Group Relative Policy Optimization (GRPO) has become a widely used approach for post-training Large Language Models (LLMs) for reasoning. In GRPO, the group gradients induced by different queries within the same mini-bat…

Rethinking Reward Signals in Video GRPO: When Scores Become Targets

2025-11-24 · Rui Li, Yuanzhi Liang, Ziqi Ni, Haibing Huang 외 arxiv

Group Relative Policy Optimization (GRPO) enables stable and preference-oriented updates via group-wise comparisons for post-training video generation. However, GRPO directly optimizes reward-induced advantages. Under su…

Video GenerationVideo Alignment

Conformal Constrained Policy Optimization for Cost-Effective LLM Agents

2025-11-14 · Wenwen Si, Sooyong Jang, Insup Lee, Osbert Bastani arxiv

While large language models (LLMs) have recently made tremendous progress towards solving challenging AI problems, they have done so at increasingly steep computational and API costs. We propose a novel strategy where we…

Multi-hop Question AnsweringReinforcement Learning