paper-with-me

홈 › Papers

Reinforcement Learning for Reasoning in Large Language Models with One Training Example

2025-04-29 · Yiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren, Lucas Liu, Baolin Peng, Hao Cheng, Xuehai He, Kuan Wang, Jianfeng Gao, Weizhu Chen, Shuohang Wang, Simon Shaolei Du, Yelong Shen

We show that reinforcement learning with verifiable reward using one training example (1-shot RLVR) is effective in incentivizing the math reasoning capabilities of large language models (LLMs). Applying RLVR to the base model Qwen2.5-Math-1.5B, we identify a single example that elevates model performance on MATH500 from 36.0% to 73.6%, and improves the average performance across six common mathematical reasoning benchmarks from 17.6% to 35.7%. This result matches the performance obtained using the 1.2k DeepScaleR subset (MATH500: 73.6%, average: 35.9%), which includes the aforementioned example. Similar substantial improvements are observed across various models (Qwen2.5-Math-7B, Llama3.2-3B-Instruct, DeepSeek-R1-Distill-Qwen-1.5B), RL algorithms (GRPO and PPO), and different math examples (many of which yield approximately 30% or greater improvement on MATH500 when employed as a single training example). In addition, we identify some interesting phenomena during 1-shot RLVR, including cross-domain generalization, increased frequency of self-reflection, and sustained test performance improvement even after the training accuracy has saturated, a phenomenon we term post-saturation generalization. Moreover, we verify that the effectiveness of 1-shot RLVR primarily arises from the policy gradient loss, distinguishing it from the "grokking" phenomenon. We also show the critical role of promoting exploration (e.g., by adding entropy loss with an appropriate coefficient) in 1-shot RLVR training. As a bonus, we observe that applying entropy loss alone, without any outcome reward, significantly enhances Qwen2.5-Math-1.5B's performance on MATH500 by 27.4%. These findings can inspire future work on RLVR data efficiency and encourage a re-examination of both recent progress and the underlying mechanisms in RLVR. Our code, model, and data are open source at https://github.com/ypwang61/One-Shot-RLVR

📄 PDF Abstract BibTeX arXiv:2504.20571

Code (1)

ypwang61/one-shot-rlvr 공식 구현 pytorch

Tasks

Domain GeneralizationMathMathematical Reasoning

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Distillation and Refinement of Reasoning in Small Language Models for Document Re-ranking

2025-04-04 · Chris Samarinas, Hamed Zamani

We present a novel approach for training small language models for reasoning-intensive document ranking that combines knowledge distillation with reinforcement learning optimization. While existing methods often rely on …

Document RankingInformation RetrievalKnowledge DistillationLanguage Modeling+4

Disentangling Reasoning and Knowledge in Medical Large Language Models

2025-05-16 · Rahul Thapa, Qingyang Wu, Kevin Wu, Harrison Zhang 외

Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reasoning with factual recall. We address thi…

DiagnosticMedQA

Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning

2025-11-15 · Yiqing Shen, Mathias Unberath arxiv

Visual reasoning may require models to interpret images and videos and respond to implicit text queries across diverse output formats, from pixel-level segmentation masks to natural language descriptions. Existing approa…

Visual Question AnsweringReinforcement LearningVisual Reasoning

The Unlearnability Phenomenon in RLVR for Language Models

2026-05-16 · Yulin Chen, He He, Chen Zhao arxiv

Reinforcement Learning with Verifiable Reward (RLVR) has proven effective in improving Large Language Model's (LLM) reasoning ability. However, the learning dynamics of RLVR remain underexplored. In this paper, we reveal…

Reinforcement LearningData Augmentation

Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

2025-03-10 · Huilin Deng, Ding Zou, Rui Ma, Hongchen Luo 외

While state-of-the-art vision-language models (VLMs) have demonstrated remarkable capabilities in complex visual-text tasks, their success heavily relies on massive model scaling, limiting their practical deployment. Sma…