paper-with-me

홈 › Papers

Reward Model Perspectives: Whose Opinions Do Reward Models Reward?

2025-10-07 · Elle arxiv

Reward models (RMs) are central to the alignment of language models (LMs). An RM often serves as a proxy for human preferences to guide downstream LM behavior. However, our understanding of RM behavior is limited. Our work (i) formalizes a framework for measuring the alignment of opinions captured by RMs, (ii) investigates the extent to which RMs demonstrate sociodemographic biases, and (iii) explores the effects of prompting to steer rewards towards the preferences of a target group. We study the subjective and diverse perspectives on controversial topics, which allows us to quantify RM perspectives in terms of their opinions, attitudes, and values. We show that RMs are poorly aligned with several demographic groups and can systematically reward harmful stereotypes, and steering alone is not enough to overcome these limitations. Our findings underscore the need for more careful consideration of RM behavior in model alignment during preference learning to prevent the propagation of unwanted social biases in the language technologies that we use.

📄 PDF Abstract BibTeX arXiv:2510.06391

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Text Simplification with Reinforcement Learning Using Supervised Rewards on Grammaticality, Meaning Preservation, and Simplicity

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Akifumi Nakamachi, Tomoyuki Kajiwara, Yuki Arase

We optimize rewards of reinforcement learning in text simplification using metrics that are highly correlated with human-perspectives. To address problems of exposure bias and loss-evaluation mismatch, text-to-text gener…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Text Generation+1

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

2025-07-28 · Shijie Zhou, Ruiyi Zhang, Huaisheng Zhu, Branislav Kveton 외 arxiv

We introduce LLaVA-Reward, an efficient reward model designed to automatically evaluate text-to-image (T2I) generations across multiple perspectives, leveraging pretrained multimodal large language models (MLLMs). Existi…

Text-to-Image Generation

Sparse Reward Subsystem in Large Language Models

2026-02-01 · Guowei Xu, Mert Yuksekgonul, James Zou arxiv

Recent studies show that LLM hidden states encode reward-related information, such as answer correctness and model confidence. However, existing approaches typically fit black-box probes on the full hidden states, offeri…

Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards

2023-06-07 · NeurIPS 2023 11 · Alexandre Ramé, Guillaume Couairon, Mustafa Shukor, Corentin Dancette 외

Foundation models are first pre-trained on vast unsupervised datasets and then fine-tuned on labeled data. Reinforcement learning, notably from human feedback (RLHF), can further align the network with the intended usage…

DiversityImage CaptioningImage GenerationText Summarization+4

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences

2025-09-03 · Arpan Mukherjee, Marcello Bullo, Deniz Gündüz arxiv

Uniform-reward reinforcement learning from human feedback (RLHF), which trains a single reward model to represent the preferences of all annotators, fails to capture the diversity of opinions across sub-populations, inad…

Reinforcement Learning