paper-with-me

홈 › Papers

F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare

2026-02-06 · Daniil Plyusov, Alexey Gorbatovski, Boris Shaposhnikov, Viacheslav Sinii, Alexey Malakhov, Daria Korotyshova, Daniil Gavrilov arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) is commonly based on group sampling to estimate advantages and stabilize policy updates. In practice, computational limits often rule out very large groups, so training proceeds with finite rollout sets that can reinforce only the correct behavior they expose. At practical group sizes, updates can miss rare-correct trajectories while still containing mixed rewards, concentrating probability on more common sampled solutions. We derive the probability of such prompt-local tail-miss events as a function of group size, showing non-monotonic behavior, and in the categorical abstraction characterize how unsampled-correct mass can shrink even as total correct mass grows. Motivated by this analysis, we propose a difficulty-aware scaling coefficient, inspired by Focal loss, that down-weights updates on high-success sampled groups. Empirically, categorical simulation illustrates the same effect in the categorical setting, Maze provides a single-solution test, and LLM experiments include a representative GRPO group-size sweep together with fixed-$N$ transfer across GRPO, DAPO, and CISPO. On Qwen2.5-7B at $N{=}8$, our method improves average math pass@256 from 64.1 $\rightarrow$ 70.3 (GRPO), 69.3 $\rightarrow$ 72.5 (DAPO), and 73.2 $\rightarrow$ 76.8 (CISPO); OOD pass@256 also improves in all three cases, without increasing group size or computational cost.

📄 PDF Abstract BibTeX arXiv:2602.06717

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Can GRPO Boost Complex Multimodal Table Understanding?

2025-09-21 · Xiaoqiang Kang, Shengen Wu, Zimu Wang, Yilin Liu 외 arxiv

Existing table understanding methods face challenges due to complex table structures and intricate logical reasoning. While supervised finetuning (SFT) dominates existing research, reinforcement learning (RL), such as Gr…

Reinforcement LearningLogical Reasoning

Spectral Policy Optimization: Coloring your Incorrect Reasoning in GRPO

2025-05-16 · Peter Chen, Xiaopeng Li, Ziniu Li, Xi Chen 외

Reinforcement learning (RL) has demonstrated significant success in enhancing reasoning capabilities in large language models (LLMs). One of the most widely used RL methods is Group Relative Policy Optimization (GRPO)~\c…

AllDiversityReinforcement Learning (RL)

MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis

2025-07-31 · Haiyun Guo, Zhiyan Hou, Yandu Sun, Jinghan He 외 arxiv

Continual instruction tuning(CIT) during the post-training phase is crucial for adapting multimodal large language models (MLLMs) to evolving real-world demands. However, the progress is hampered by the lack of benchmark…

Continual Learning

Joint Continual Learning of Local Language Models and Cloud Offloading Decisions with Budget Constraints

2026-01-29 · Evan Chen, Wenzhi Fang, Shiqiang Wang, Christopher Brinton arxiv

Locally deployed Small Language Models (SLMs) must continually support diverse tasks under strict memory and computation constraints, making selective reliance on cloud Large Language Models (LLMs) unavoidable. Regulatin…

Reinforcement LearningMathematical ReasoningContinual LearningCode Generation

Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training

2026-04-02 · William Hoy, Binxu Wang, Xu Pan arxiv

Evolution Strategies (ES) have emerged as a scalable gradient-free alternative to reinforcement learning based LLM fine-tuning, but it remains unclear whether comparable task performance implies comparable solutions in p…

Reinforcement Learning