paper-with-me

Papers

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration

2026-06-25 · Arnav Raj arxiv

Reward models for Reinforcement Learning from Human Feedback (RLHF) pool preferences across thousands of annotators and fit one global affine calibrator, collapsing raters with systematically different rating-scale offsets and slopes into a single average-rater fit that does not match any individual annotator. PEBS is a per-rater empirical-Bayes shrinkage estimator: it fits per-rater affine calibrators on a held-out slice of each annotator's ratings and applies Morris-James-Stein empirical-Bayes shrinkage toward the population mean, in closed form and without retraining the reward model. On PRISM, PEBS reduces within-user held-out RMSE by 8.58% over the pooled population-slope baseline. The procedure replicates on PluriHarms harm ratings (Qwen-2.5 base, in-family) with a +9.66% RMSE reduction over the same population-slope baseline. PEBS is a closed-form post-hoc estimator for annotator-specific affine calibration in RLHF reward modeling; it leaves the reward base model unchanged and estimates only the rater-level map used at inference time for new ratings.

📄 PDF Abstract BibTeX arXiv:2606.27578

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Rater State Bias in RLHF Preference Data: An Audit Framework

2026-04-14 · Elena Kopteva, Vitaliy Hlynianyi-Zhuk arxiv

We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater's state during annotati…

Online Bandit Learning with Offline Preference Data for Improved RLHF

2024-06-13 · Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng Wen

Reinforcement Learning with Human Feedback (RLHF) is at the core of fine-tuning methods for generative AI models for language and images. Such feedback is often sought as rank or preference feedback from human raters, as…

Active Learning

Empirical Bayes shrinkage (mostly) does not correct the measurement error in regression

2025-03-24 · Jiafeng Chen, Jiaying Gu, Soonwoo Kwon

In the value-added literature, it is often claimed that regressing on empirical Bayes shrinkage estimates corrects for the measurement error problem in linear regression. We clarify the conditions needed; we argue that t…

Attribute

Flexible shrinkage in high-dimensional Bayesian spatial autoregressive models

2018-05-28

This article introduces two absolutely continuous global-local shrinkage priors to enable stochastic variable selection in the context of high-dimensional matrix exponential spatial specifications. Existing approaches as…

Variable SelectionVocal Bursts Intensity Prediction

Forecasting macroeconomic data with Bayesian VARs: Sparse or dense? It depends!

2022-06-10 · Luis Gruber, Gregor Kastner

Vector autogressions (VARs) are widely applied when it comes to modeling and forecasting macroeconomic variables. In high dimensions, however, they are prone to overfitting. Bayesian methods, more concretely shrinkage pr…

Variable Selection