STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment
Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each preference dimension enter policy optimization? We propose \methodname, a stability-guided active-set controller for controlled objective admission. \methodname starts from a small active set, retains admitted objectives, and expands when reward-deviation gates indicate low recent deviation or a patience budget is exhausted. A probing phase estimates a hard-to-easy order, and adaptive weighting emphasizes underperforming active dimensions. Automatic evaluations with 15 training preferences and 16 held-out benchmark columns show that \methodname obtains higher averages than simultaneous scalarization and shared-budget adapted baselines. Component ablations and expansion dynamics further support cumulative retention, gated admission, and probing-derived ordering as useful design choices in this setting. These results position objective-entry timing as a concrete control variable in reward-vector RLHF.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Personalized Generation
Personalized generation systems retrieve user history by request--memory relevance and inject it into the model context. Yet relevant history may concern the wrong preference aspect, duplicate public information, or prov…
Semantic RetrievalDecentralizing Centralized Matching Markets: Implications from Early Offers in University Admissions
The matching literature often recommends market centralization under the assumption that agents know their own preferences and that their preferences are fixed. We find counterevidence to this assumption in a quasi-exper…
Retention Consequence in Lifecycle Memory Control
Persistent memory can fail after successful admission: a premise is written, then becomes a silent assumption, and later maintenance treats it as ordinary residue to be compressed, demoted, or evicted. We study this post…
FedQueue: Queue-Aware Federated Learning for Cross-Facility HPC Training
Federated learning (FL) across multiple HPC facilities faces stochastic admission delays from batch schedulers that dominate wall-clock time. Synchronous FL suffers from severe stragglers, while asynchronous FL accumulat…
Federated LearningWhy Do Students Lie and Should We Worry? An Analysis of Non-truthful Reporting
A core aspect in market design is to encourage participants to truthfully report their preferences to ensure efficiency and fairness. Our research paper analyzes the factors that contribute to and the consequences of stu…
Fairness