paper-with-me

홈 › Papers

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment

2026-05-17 · Yucong Huang, Xiucheng Li, Kaiqi Zhao, Jing Li arxiv

Standard RLHF relies on transitive scalar rewards, failing to capture the cyclic nature of human preferences. While some approaches like the General Preference Model (GPM) address this, we identify a theoretical limitation: their implicit formulation entangles hierarchy with cyclicity, failing to guarantee dominant solutions. To address this, we propose the Hybrid Reward-Cyclic (HRC) model, which utilizes game-theoretic decomposition to explicitly disentangle preferences into orthogonal transitive (scalar) and cyclic (vector) components. Complementing this, we introduce Dynamic Self-Play Preference Optimization (DSPPO), which treats alignment as a time-varying game to progressively guide the policy toward the Nash equilibrium. Synthetic data experiments further validate HRC's structural superiority in mixed transitive--cyclic settings, where HRC converges faster and achieves higher accuracy than GPM. Experiments on RewardBench 2 demonstrate that HRC consistently improves over both BT and GPM baselines (e.g., +1.23% on Gemma-2B-it). In particular, its superior performance in the Ties domain empirically validates the model's robustness in handling complex, non-strict preferences. Extensive downstream evaluations on AlpacaEval 2.0, Arena-Hard-v0.1, and MT-Bench confirm the efficacy of our framework. Notably, when using Gemma-2B-it as the base preference model, HRC+DSPPO achieves a peak length-controlled win-rate of 44.75% on AlpacaEval 2.0 and 46.8% on Arena-Hard-v0.1, significantly outperforming SPPO baselines trained with BT or GPM. Our code is publicly available at https://github.com/lab-klc/Hybrid-Reward-Cyclic.

📄 PDF Abstract BibTeX arXiv:2605.17342

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Escaping Arrow's Theorem: The Advantage-Standard Model

2021-08-02 · Wesley H. Holliday, Mikayla Kelley

There is an extensive literature in social choice theory studying the consequences of weakening the assumptions of Arrow's Impossibility Theorem. Much of this literature suggests that there is no escape from Arrow-style …

model

Similarity Suppresses Cyclicity: Why Similar Competitors Form Hierarchies

2022-05-16 · Christopher Cebra, Alexander Strang

Competitive systems can exhibit both hierarchical (transitive) and cyclic (intransitive) structures. Despite theoretical interest in cyclic competition, which offers richer dynamics, and occupies a larger subset of the s…

Form

Transitivity, Time Consumption, and Quality of Preference Judgments in Crowdsourcing

2021-04-18 · Kai Hui, Klaus Berberich

Preference judgments have been demonstrated as a better alternative to graded judgments to assess the relevance of documents relative to queries. Existing work has verified transitivity among preference judgments when co…

Describing Sen's Transitivity Condition in Inequalities and Equations

2022-04-08 · Fujun Hou

In social choice theory, Sen's value restriction condition is a sufficiency condition restricted to individuals' ordinal preferences so as to obtain a transitive social preference under the majority decision rule. In thi…

Position

Intransitivity in Theory and in the Real World

2015-07-11

This work considers reasons for and implications of discarding the assumption of transitivity, which (transitivity) is the fundamental postulate in the utility theory of Von Neumann and Morgenstern, the adiabatic accessi…