paper-with-me

홈 › Papers

Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order

2025-12-03 · Prakhar Gupta, Vaibhav Gupta arxiv

Post-training with reinforcement learning (RL) typically optimizes a single scalar objective and ignores structure in how solutions are produced. We ask whether a scalar hint toward a canonical solver ordering, used only during RL post-training, improves performance even when fine-tuned on randomized solution sequences. On Zebra puzzles, we fine-tune a Transformer on randomized solution orders, then post-train it with Group Relative Policy Optimization (GRPO) using two rewards: a sparse task reward that is 1 only when the puzzle is fully solved, and an ordering reward that increases when the model's emission order aligns with the canonical solver order. To compare signals cleanly, we combine them via fixed mixtures and use a simple bootstrapped scaling to equalize component magnitudes at initialization. Mixed rewards generally outperform task-only optimization, suggesting that coarse ordering signals can steer RL post-training toward canonical trajectories without modifying supervised data or architecture.

📄 PDF Abstract BibTeX arXiv:2512.04277

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

EditGRPO: Reinforcement Learning with Post-Rollout Edits for Clinically Accurate Chest X-Ray Report Generation

2025-09-26 · Kai Zhang, Christopher Malon, Lichao Sun, Martin Renqiang Min arxiv

Radiology report generation requires advanced medical image analysis, effective temporal reasoning, and accurate text generation. Although recent innovations, particularly multimodal large language models, have shown imp…

Reinforcement LearningDomain GeneralizationText Generation

Improving Unsupervised Visual Program Inference with Code Rewriting Families

2023-09-26 · ICCV 2023 1 · Aditya Ganeshan, R. Kenny Jones, Daniel Ritchie

Programs offer compactness and structure that makes them an attractive representation for visual data. We explore how code rewriting can be used to improve systems for inferring programs from visual data. We first propos…

Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards

2025-05-30 · Ruipeng Jia, Yunyi Yang, Yongbo Gai, Kai Luo 외

Reinforcement learning with verifiable rewards (RLVR) has enabled large language models (LLMs) to achieve remarkable breakthroughs in reasoning tasks with objective ground-truth answers, such as mathematics and code gene…

Code Generation

Meta-Gradient Reinforcement Learning

2018-05-24 · NeurIPS 2018 12 · Zhongwen Xu, Hado van Hasselt, David Silver

The goal of reinforcement learning algorithms is to estimate and/or optimise the value function. However, unlike supervised learning, no teacher or oracle is available to provide the true value function. Instead, the maj…

Meta-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Quantifying Empirical Compute-Supervision Tradeoffs in RLVR

2026-05-24 · Ryo Mitsuhashi, Patrick Chen, Isabelle Tseng, Jasin Cekinmez 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training language models, but in practice, verifiers are rarely perfect. Recent theoretical work predicts that verifier noise …

Reinforcement Learning