paper-with-me

홈 › Papers

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance

2026-05-14 · Kai Yan, Alexander G. Schwing, Yu-Xiong Wang arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has achieved great success in developing Large Language Models (LLMs) with chain-of-thought rollouts for many tasks such as math and coding. Nevertheless, RLVR struggles with sample efficiency on difficult problems where correct rollouts are hard to generate. Prior works propose to address this issue via demonstration-guided RLVR, i.e., to conduct Supervised FineTuning (SFT) when RL fails; however, SFT often requires a lot of data, which can be expensive to acquire. In this paper, we propose FEST, a FEw-ShoT demonstration-guided RLVR algorithm. It attains compelling results with only 128 demonstrations randomly selected from an SFT dataset. We find that three components are vital for the success: supervised signal, on-policy signal, and decaying weights on the few-shot SFT dataset to prevent overfitting from multiple-epoch training. On several benchmarks, FEST outperforms baselines with magnitudes less SFT data, even matching their performance with full dataset.

📄 PDF Abstract BibTeX arXiv:2605.15012

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

2026-05-25 · Li Wang, Xiaodong Lu, Xiaohan Wang, Yikun Ban 외 arxiv

Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR intrinsically relies on ground-truth labe…

Reinforcement Learning

The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR

2026-02-02 · Israel Adewuyi, Solomon Okibe, Vladmir Ivanov arxiv

The Lottery Ticket Hypothesis demonstrated that sparse subnetworks can match full-model performance, suggesting parameter redundancy. Meanwhile, in Reinforcement Learning with Verifiable Rewards (RLVR), recent work has s…

Reinforcement Learning

From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation

2026-01-26 · Yuxin Jiang, Yufei Wang, Qiyuan Zhang, Xingshan Zeng 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) succeeds in reasoning tasks (e.g., math and code) by checking the final verifiable answer (i.e., a verifiable dot signal). However, extending this paradigm to open-en…

Reinforcement Learning

LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards

2026-03-02 · Guanzheng Chen, Michael Qizhe Shieh, Lidong Bing arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Large Language Models (LLMs) by optimizing them against factual outcomes. However, this paradigm falters in l…

Reinforcement Learning

Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models

2025-09-29 · Yuhua Jiang, Jiawei Huang, Yufeng Yuan, Xin Mao 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for enhancing Large Language Models (LLMs) on complex reasoning tasks. However, existing methods suffer from an exploration dilemma: the sharply …

Reinforcement LearningMathematical Reasoning