paper-with-me

홈 › Papers

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting

2025-08-08 · Si Shen, Peijun Shen, Wenhua Zhao, Danhao Zhu arxiv

Group-Relative Policy Optimization (GRPO) is a key technique for training large reasoning models, yet it suffers from a critical vulnerability: the \emph{Think-Answer Mismatch}, where noisy reward signals corrupt the learning process. This problem is most severe in unbalanced response groups, paradoxically degrading the signal precisely when it should be most informative. To address this challenge, we propose Stable Group-Relative Policy Optimization (S-GRPO), a principled enhancement that derives optimal, noise-aware advantage weights to stabilize training. Our comprehensive experiments on mathematical reasoning benchmarks demonstrate S-GRPO's effectiveness and robustness. On various models, S-GRPO significantly outperforms DR. GRPO, achieving performance gains of +2.5% on Qwen-Math-7B-Base, +2.2% on Llama-3.2-3B-Base, and +2.4% on Qwen-Math-1.5B-Instruct. Most critically, while standard GRPO fails to learn under 20% synthetic reward noise, S-GRPO maintains stable learning progress. These results highlight S-GRPO's potential for more robust and effective training of large-scale reasoning models. \footnote{Code and data are available at: https://github.com/shenpeijun0212/S-GRPO

📄 PDF Abstract BibTeX arXiv:2508.05928

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment

2026-06-12 · Jiayue Cao, Zhicong Lu, Xuehan Sun, Wei Jia 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has successfully elicited the reasoning capabilities of large language models, motivating its extension to multimodal scenarios. Existing methods primarily focus on i…

Reinforcement LearningMultimodal Reasoning

Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

2026-02-16 · Guangyue Peng, Zongchao Chen, Wen Luo, Yuntao Wen 외 arxiv

Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but it risks producing post-hoc rationalizations: when models can see the answer during generation, a systematic train-infer…

Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization

2026-07-07 · Kaishen Wang, Tong Zheng, Xuehao Cui, Ruibo Chen 외 arxiv

Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such thinking often improves overall performa…

Reinforcement LearningQuestion Answering

When Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy

2025-05-28 · Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández 외

Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as a…

Explore Briefly, Then Decide: Mitigating LLM Overthinking via Cumulative Entropy Regulation

2025-10-02 · Yi Bin, Tianyi Jiang, Yujuan Ding, Kainian Zhu 외 arxiv

Large Language Models (LLMs) have demonstrated remarkable reasoning abilities on complex problems using long Chain-of-Thought (CoT) reasoning. However, they often suffer from overthinking, meaning generating unnecessaril…