paper-with-me

홈 › Papers

Spectral Policy Optimization: Coloring your Incorrect Reasoning in GRPO

2025-05-16 · Peter Chen, Xiaopeng Li, Ziniu Li, Xi Chen, Tianyi Lin

Reinforcement learning (RL) has demonstrated significant success in enhancing reasoning capabilities in large language models (LLMs). One of the most widely used RL methods is Group Relative Policy Optimization (GRPO)~\cite{Shao-2024-Deepseekmath}, known for its memory efficiency and success in training DeepSeek-R1~\cite{Guo-2025-Deepseek}. However, GRPO stalls when all sampled responses in a group are incorrect -- referred to as an \emph{all-negative-sample} group -- as it fails to update the policy, hindering learning progress. The contributions of this paper are two-fold. First, we propose a simple yet effective framework that introduces response diversity within all-negative-sample groups in GRPO using AI feedback. We also provide a theoretical analysis, via a stylized model, showing how this diversification improves learning dynamics. Second, we empirically validate our approach, showing the improved performance across various model sizes (7B, 14B, 32B) in both offline and online learning settings with 10 benchmarks, including base and distilled variants. Our findings highlight that learning from all-negative-sample groups is not only feasible but beneficial, advancing recent insights from \citet{Xiong-2025-Minimalist}.

📄 PDF Abstract BibTeX arXiv:2505.11595

Code (0)

등록된 구현이 없습니다.

Tasks

AllDiversityReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution

2026-05-22 · Hongbo Wang, Huaibo Huang, Pin Wang, Jinhua Hao 외 arxiv

Generative priors in Image Super-Resolution (SR) often compromise faithful restoration, we attribute this limitation to a fundamental spectral misalignment between isotropic objectives and the intrinsic natural image man…

Image Super-Resolution

Weisfeiler and Leman Go Infinite: Spectral and Combinatorial Pre-Colorings

2022-01-31 · Or Feldman, Amit Boyarski, Shai Feldman, Dani Kogan 외

Graph isomorphism testing is usually approached via the comparison of graph invariants. Two popular alternatives that offer a good trade-off between expressive power and computational efficiency are combinatorial (i.e., …

Computational EfficiencyIsomorphism TestingOpen-Ended Question Answering

Learning Combinatorial Node Labeling Algorithms

2021-06-07 · Lukas Gianinazzi, Maximilian Fries, Nikoli Dryden, Tal Ben-Nun 외

We present a novel neural architecture to solve graph optimization problems where the solution consists of arbitrary node labels, allowing us to solve hard problems like graph coloring. We train our model using reinforce…

BIG-bench Machine LearningGraph AttentionReinforcement Learning (RL)

RecolorCloud: A Point Cloud Tool for Recoloring, Segmentation, and Conversion

2023-10-19 · Esteban Segarra Martinez, Ryan P. McMahan

Point clouds are a 3D space representation of an environment that was recorded with a high precision laser scanner. These scanners can suffer from environmental interference such as surface shading, texturing, and reflec…

Semantic Segmentation

Estimation of blood oxygenation with learned spectral decoloring for quantitative photoacoustic imaging (LSD-qPAI)

2019-02-15 · Janek Gröhl, Thomas Kirchner, Tim Adler, Lena Maier-Hein

One of the main applications of photoacoustic (PA) imaging is the recovery of functional tissue properties, such as blood oxygenation (sO2). This is typically achieved by linear spectral unmixing of relevant chromophores…