paper-with-me

홈 › Papers

Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

2025-08-13 · Vaishnavi Shrivastava, Ahmed Awadallah, Vidhisha Balachandran, Shivam Garg, Harkirat Behl, Dimitris Papailiopoulos arxiv

Large language models trained with reinforcement learning with verifiable rewards tend to trade accuracy for length--inflating response lengths to achieve gains in accuracy. While longer answers may be warranted for harder problems, many tokens are merely "filler": repetitive, verbose text that makes no real progress. We introduce GFPO (Group Filtered Policy Optimization), which curbs this length explosion by sampling larger groups per problem during training and filtering responses to train on based on two key metrics: (1) response length and (2) token efficiency: reward per token ratio. By sampling more at training time, we teach models to think less at inference time. On the Phi-4-reasoning model, GFPO cuts GRPO's length inflation by 46-71% across challenging STEM and coding benchmarks (AIME 24/25, GPQA, Omni-MATH, LiveCodeBench) while maintaining accuracy. Optimizing for reward per token further increases reductions in length inflation to 71-85%. We also propose Adaptive Difficulty GFPO, which dynamically allocates more training resources to harder problems based on real-time difficulty estimates, improving the balance between computational efficiency and accuracy especially on difficult questions. GFPO demonstrates that increased training-time compute directly translates to reduced test-time compute--a simple yet effective trade-off for efficient reasoning.

📄 PDF Abstract BibTeX arXiv:2508.09726

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyReinforcement Learning

Similar Papers 제목 키워드 기반

Deformable Wiener Filter for Future Video Coding

2026-06-01 · Xuewei Meng, Chuanmin Jia, Xinfeng Zhang, Shanshe Wang 외 arxiv

In-loop filters have attracted increasing attention due to the remarkable noise-reduction capability in the hybrid video coding framework. However, the existing in-loop filters in Versatile Video Coding (VVC) mainly take…

FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering

2025-01-13 · Erik Henriksson, Otto Tarkka, Filip Ginter

Data quality is crucial for training Large Language Models (LLMs). Traditional heuristic filters often miss low-quality text or mistakenly remove valuable content. In this paper, we introduce an LLM-based line-level filt…

DescriptiveHellaSwag

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

2026-01-30 · Seanie Lee, Sangwoo Park, Yumin Choi, Gyeongman Kim 외 arxiv

Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-thought (CoT) reasoning. However, this over-optimization often prioritiz…

Reinforcement Learning

Thinkless: LLM Learns When to Think

2025-05-19 · Gongfan Fang, Xinyin Ma, Xinchao Wang

Reasoning Language Models, capable of extended chain-of-thought reasoning, have demonstrated remarkable performance on tasks requiring complex logical inference. However, applying elaborate reasoning for all queries ofte…

GSM8KMath

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs

2025-07-03 · Purbesh Mitra, Sennur Ulukus arxiv

Recent advancements in the reasoning capabilities of large language models (LLMs) show that employing group relative policy optimization (GRPO) algorithm for reinforcement learning (RL) training allows the models to use …

Reinforcement Learning