paper-with-me

홈 › Papers

What should post-training optimize? A test-time scaling law perspective

2026-05-11 · Muheng Li, Jian Qian, Wenlong Mou arxiv

Large language models are increasingly deployed with test-time strategies: sample $N$ responses, score them with a reward model or verifier, and return the best. This deployment rule exposes a mismatch in post-training: standard objectives optimize the mean reward of a single response, whereas best-of-$N$ performance is governed by the upper tail of the reward distribution. Recent test-time-aware objectives partly address this mismatch, but typically assume that training can use the same per-prompt rollout budget as deployment, which is impractical when post-training must cover many prompts while deployment can allocate much larger per-prompt test-time compute. We study this budget-mismatch regime, where only $m\ll N$ per-prompt rollouts are available during training but the target objective is best-of-$N$ deployment. Under structural assumptions on the reward tails, we show that the policy gradient of the best-of-$N$ objective can be approximated from a much smaller rollout group by extrapolating upper-tail statistics. This yields a family of Tail-Extrapolated estimators for best-of-$N$-oriented post-training: a simple direct estimator, Tail-Extrapolated Advantage (TEA), and a fixed-order debiased Prefix-TEA estimator based on moment cancellation. Experiments on instruction-following tasks show that TEA and Prefix-TEA improve best-of-$N$ performance across different language models, reward models and datasets under various training and test-time budget settings.

📄 PDF Abstract BibTeX arXiv:2605.10716

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning

2020-07-06 · ICML 2020 1 · Silviu Pitis, Harris Chan, Stephen Zhao, Bradly Stadie 외

What goals should a multi-goal reinforcement learning agent pursue during training in long-horizon tasks? When the desired (test time) goal distribution is too distant to offer a useful learning signal, we argue that the…

Multi-Goal Reinforcement Learningreinforcement-learningReinforcement Learning (RL)

Data Excellence for AI: Why Should You Care

2021-11-19 · Lora Aroyo, Matthew Lease, Praveen Paritosh, Mike Schaekermann

The efficacy of machine learning (ML) models depends on both algorithms and data. Training data defines what we want our models to learn, and testing data provides the means by which their empirical progress is measured.…

Inverse Reward Design

2017-11-08 · NeurIPS 2017 12 · Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart Russell 외

Autonomous agents optimize the reward function we give them. What they don't know is how hard it is for us to design a reward function that actually captures what we want. When designing the reward, we might think of som…

TEST_POSITIVE at W-NUT 2020 Shared Task-3: Cross-task modeling

2020-11-01 · EMNLP (WNUT) 2020 11 · Chacha Chen, Chieh-Yang Huang, Yaqi Hou, Yang Shi 외

The competition of extracting COVID-19 events from Twitter is to develop systems that can automatically extract related events from tweets. The built system should identify different pre-defined slots for each event, in …

Extracting COVID-19 Events from TwitterLanguage ModelingLanguage ModellingMulti-Task Learning+4

The Blocker Postulates for Measures of Voting Power

2022-05-17 · Arash Abizadeh, Adrian Vetta

A proposed measure of voting power should satisfy two conditions to be plausible: first, it must be conceptually justified, capturing the intuitive meaning of what voting power is; second, it must satisfy reasonable post…