paper-with-me

Papers

Hidden Consensus:Preference-Validity Compression in Human Feedback

2026-06-09 · Dorcas Chia Ern Chua, Karen Myn Hui Lee, Jia Yue Tan, Zhen Xue Gue, Norzalena Abdul Hamid, Azima Binti Azmi, Keat Mei Yeong, Aizat Izyani binti Mujab, Hafsah Noor Azam, Chee Guo Khoo, Han Ying Lim, Chee Seng Chan arxiv

Standard RLHF pipelines often reduce heterogeneous human judgments into a single scalar reward target. We argue that this reduction can mis-measure alignment in structurally plural societies, where disagreement may reflect culturally, historically, linguistically, regionally, or normatively grounded interpretations rather than annotation noise. We call this failure Preference-Validity Compression, the collapse of multiple plural-valid response options into a single optimization target. Using Malaysia as a diagnostic setting, we analyze RLHF-style feedback aggregation through preference events linking prompts, responses, and acceptability judgments across interpretive frames. Across 321 preference events from 20 participants and 107 trio-annotated prompts, 79% of prompts contain more than one majority-supported response that single-winner aggregation would discard, and apparent dominance gaps between top responses diminish when all majority-supported options are considered. Participants frequently select multiple acceptable responses, and discarded responses demonstrably reflect coherent local, practical, or cultural frames. These findings show that majority aggregation in this corpus measures argmax acceptability rather than plural alignment. We treat this as a measurement-validity issue and argue that future alignment methods should satisfy Validity-Preserving Consistency, remaining stable across plural-valid interpretive frames rather than collapsing them into a single reward target.

📄 PDF Abstract BibTeX arXiv:2606.10569

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating Alignment of Behavioral Dispositions in LLMs

2026-02-11 · Amir Taubenfeld, Zorik Gekhman, Lior Nezry, Omri Feldman 외 arxiv

As LLMs integrate into our daily lives, understanding their behavior becomes essential. In this work, we focus on behavioral dispositions$-$the underlying tendencies that shape responses in social contexts$-$and introduc…

Fine-tuning language models to find agreement among humans with diverse preferences

2022-11-28 · Michiel A. Bakker, Martin J. Chadwick, Hannah R. Sheahan, Michael Henry Tessler 외

Recent work in large language modeling (LLMs) has used fine-tuning to align outputs with the preferences of a prototypical user. This work assumes that human preferences are static and homogeneous across individuals, so …

Language ModelingLanguage Modelling

AlignGroup: Learning and Aligning Group Consensus with Member Preferences for Group Recommendation

2024-09-04 · Jinfeng Xu, Zheyu Chen, Jinze Li, Shuo Yang 외

Group activities are important behaviors in human society, providing personalized recommendations for groups is referred to as the group recommendation task. Existing methods can usually be categorized into two strategie…

Decision Making

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization

2026-06-30 · Ke Zhang, Patricio Gallardo Candela, Sudhir Murthy, Yi Xie 외 arxiv

Theorem-proving benchmarks evaluate proof search against fixed formal statements, but natural-language-to-Lean formalization must generate the formal statement itself. In this setting, compilation is only a validity chec…

Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders

2026-07-11 · Cedric Waterschoot, Nava Tintarev, Francesco Barile arxiv

Previous work in group recommender systems has demonstrated a sensitivity to the distribution of preferences within a group. Specifically, the selection of the preference aggregation strategy benefits from considering su…