paper-with-me

홈 › Papers

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime

2026-05-06 · Tianshu Zhu, Wenyu Zhang, Xiaoying Zuo, Lun Tian, Haotian Zhao, Yucheng Zeng, Jingnan Gu, Daxiang Dong, Jianmin Wu, Dawei Yin, Dou Shen arxiv

Agentic reinforcement learning (RL) for software engineering spends much of its compute on stateful trajectories whose grouped binary rewards are highly skewed and weakly contrastive. We frame this as pass-rate control and show that the binary reward-side signal is strongest near a 50% rollout pass rate under four criteria: reward entropy, group-filtering survival, leave-one-out (RLOO) advantage energy under Group Relative Policy Optimization (GRPO), and success-failure pair count. We propose Prefix Sampling (PS), which replays self-generated trajectory prefixes to steer skewed groups toward this regime: successful prefixes give mostly failing groups a head start, while failing prefixes handicap mostly passing groups. Replayed states are reconstructed through the existing rollout path, and replayed tokens are masked from the loss so optimization applies only to current-policy continuations. On SWE-bench Verified, PS reaches the baseline high-score regime within evaluation variability while delivering 2.01x and 1.55x end-to-end wall-clock speedups on Qwen3-14B and Qwen3-32B; the 14B peak improves from 0.274 to 0.295. AIME 2025 experiments on 4B and 8B show the same pass-rate-control pattern, and 4B ablations attribute gains to replay, bidirectional coverage, and adaptive control.

📄 PDF Abstract BibTeX arXiv:2605.05112

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Leverage Is Not Reach: A Control-Window Law for Single-Neuron Steering in Language Models

2026-06-18 · Hongliang Liu arxiv

Aligned language models gate behaviors such as refusal and language routing through sparse feed forward neurons, yet no theory predicts when a single neuron intervention controls a behavior coherently rather than collaps…

ROAST: Rollout-based On-distribution Activation Steering Technique

2026-02-15 · Xuanbo Su, Hao Luo, Yingfang Zhang, Lijun Zhang arxiv

Activation steering provides parameter-efficient control over large language models (LLMs) at inference time, but many methods rely on off-distribution supervision and discrete masking, leading to brittle interventions. …

VSPO: Vector-Steered Policy Optimization for Behavioral Control

2026-05-15 · Xuechen Zhang, Zijian Huang, Kai Yang, Weijia Zhang 외 arxiv

Modern language models often need to optimize a primary accuracy objective while also accommodating secondary behavioral preferences, such as verbosity, agreeableness, or the level of technical expertise in its response.…

When is Your LLM Steerable?

2026-06-10 · Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi 외 arxiv

Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration. Findin…

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

2026-07-16 · Jihoon Hong, Julian Skifstad, Qiyue Dai, Alice Chan 외 arxiv

World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations a…