paper-with-me

홈 › Papers

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

2025-10-30 · Rajiv Movva, Smitha Milli, Sewon Min, Emma Pierson arxiv

Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback data encodes. While prior work studies preferences over certain attributes (e.g., length or sycophancy), automatically extracting relevant features without pre-specifying hypotheses remains challenging. We introduce What's In My Human Feedback? (WIMHF), a method to explain feedback data using sparse autoencoders. WIMHF characterizes both (1) the preferences a dataset is capable of measuring and (2) the preferences that the annotators actually express. Across 7 datasets, WIMHF identifies a small number of human-interpretable features that account for the majority of the preference prediction signal achieved by black-box models. These features reveal a wide diversity in what humans prefer, and the role of dataset-level context: for example, users on Reddit prefer informality and jokes, while annotators in HH-RLHF and PRISM disprefer them. WIMHF also surfaces potentially unsafe preferences, such as that LMArena users tend to vote against refusals, often in favor of toxic content. The learned features enable effective data curation: re-labeling the harmful examples in Arena yields large safety gains (+37%) with no cost to general performance. They also allow fine-grained personalization: on the Community Alignment dataset, we learn annotator-specific weights over subjective features that improve preference prediction. WIMHF provides a human-centered analysis method for practitioners to better understand and use preference data.

📄 PDF Abstract BibTeX arXiv:2510.26202

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Is Missing: Interpretable Ratings for Large Language Model Outputs

2026-02-17 · Nicholas Stranges, Yimin Yang arxiv

Current Large Language Model (LLM) preference learning methods such as Proximal Policy Optimization and Direct Preference Optimization learn from direct rankings or numerical ratings of model outputs, these rankings are …

Can a Unimodal Language Agent Provide Preferences to Tune a Multimodal Vision-Language Model?

2026-01-10 · Sazia Tabasum Mim, Jack Morris, Manish Dhakal, Yanming Xiu 외 arxiv

To explore a more scalable path for adding multimodal capabilities to existing LLMs, this paper addresses a fundamental question: Can a unimodal LLM, relying solely on text, reason about its own informational needs and p…

Text Generation

Learning Interpretable Models of Aircraft Handling Behaviour by Reinforcement Learning from Human Feedback

2023-05-26 · Tom Bewley, Jonathan Lawry, Arthur Richards

We propose a method to capture the handling abilities of fast jet pilots in a software model via reinforcement learning (RL) from human preference feedback. We use pairwise preferences over simulated flight trajectories …

Reinforcement Learning (RL)

WikiPersonas: What Can We Learn From Personalized Alignment to Famous People?

2025-05-19 · Zilu Tang, Afra Feyza Akyürek, Ekin Akyürek, Derry Wijaya

Preference alignment has become a standard pipeline in finetuning models to follow \emph{generic} human preferences. Majority of work seeks to optimize model to produce responses that would be preferable \emph{on average…

Nonverbal Robot Feedback for Human Teachers

2019-11-06 · Sandy H. Huang, Isabella Huang, Ravi Pandya, Anca D. Dragan

Robots can learn preferences from human demonstrations, but their success depends on how informative these demonstrations are. Being informative is unfortunately very challenging, because during teaching, people typicall…