paper-with-me

홈 › Papers

Learn Hard Problems During RL with Reference Guided Fine-tuning

2026-03-01 · Yangzhen Wu, Shanda Li, Zixin Wen, Xin Zhou, Ameet Talwalkar, Yiming Yang, Wenhao Huang, Tianle Cai arxiv

Reinforcement learning (RL) for mathematical reasoning can suffer from reward sparsity: for challenging problems, LLM fails to sample any correct trajectories, preventing RL from receiving meaningful positive feedback. At the same time, there often exist human-written reference solutions along with the problem (e.g., problems from AoPS), but directly fine-tuning on these solutions offers no benefit because models often cannot imitate human proofs that lie outside their own reasoning distribution. We introduce Reference-Guided Fine-Tuning (ReGFT), a simple and effective method that utilizes human-written reference solutions to synthesize positive trajectories on hard problems and train on them before RL. For each problem, we provide the model with a partial reference solution and let it generate its own reasoning trace, ensuring the resulting trajectories remain in the model's reasoning space while still benefiting from reference guidance. Fine-tuning on these reference-guided trajectories increases the number of solvable problems and produces a checkpoint that receives more positive rewards during RL. Across three benchmarks (AIME24, AIME25, BeyondAIME), ReGFT consistently improves supervised accuracy, accelerates DAPO training, and raises the final performance plateau of RL. Our results show that ReGFT effectively overcomes reward sparsity and unlocks stronger RL-based mathematical reasoning.

📄 PDF Abstract BibTeX arXiv:2603.01223

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration

2026-01-26 · Yuxiao Qu, Amrith Setlur, Virginia Smith, Ruslan Salakhutdinov 외 arxiv

Reinforcement learning (RL) has improved the reasoning abilities of large language models (LLMs), yet state-of-the-art methods still fail to learn on many training problems. On hard problems, on-policy RL rarely explores…

Reinforcement Learning

Demonstration-Guided Humanoid Stand-Up on an Emulated Deformable Surface

2026-08-21 · Aniruddh Kushwah, Vyankatesh Ashtekar, Ashish Dutta arxiv

This paper presents a reference-guided reinforcement learning framework to generate stand-up motion for a 29-DOF Unitree G1 humanoid on deformable soft ground, using a human demonstration recorded on hard ground. The ter…

Reinforcement Learning

GASP: Guided Asymmetric Self-Play For Coding LLMs

2026-03-16 · Swadesh Jana, Cansu Sancaktar, Tomáš Daniš, Georg Martius 외 arxiv

Asymmetric self-play has emerged as a promising paradigm for post-training large language models, where a teacher continually generates questions for a student to solve at the edge of the student's learnability. Although…

References Improve LLM Alignment in Non-Verifiable Domains

2026-02-18 · Kejian Shi, Yixin Liu, Peifeng Wang, Alexander R. Fabbri 외 arxiv

While Reinforcement Learning with Verifiable Rewards (RLVR) has shown strong effectiveness in reasoning tasks, it cannot be directly applied to non-verifiable domains lacking ground-truth verifiers, such as LLM alignment…

Reinforcement Learning

Preference-Guided Planning: An Active Elicitation Approach

2018-04-19 · Mayukh Das, Phillip Odom, Md. Rakibul Islam, Janardhan Rao 외

Planning with preferences has been employed extensively to quickly generate high-quality plans. However, it may be difficult for the human expert to supply this information without knowledge of the reasoning employed by …