paper-with-me

홈 › Papers

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

2025-06-14 · Sara Rajaram, R. James Cotton, Fabian H. Sinz arxiv

Preference-based Reinforcement Learning (PbRL) entails a variety of approaches for aligning models with human intent to alleviate the burden of reward engineering. However, most previous PbRL work has not investigated the robustness to labeler errors, inevitable with labelers who are non-experts or operate under time constraints. We introduce Similarity as Reward Alignment (SARA), a simple contrastive framework that is both resilient to noisy labels and adaptable to diverse feedback formats. SARA learns a latent representation of preferred samples and computes rewards as similarities to the learned latent. On preference data with varying realistic noise rates, we demonstrate competitive and more stable performance on continuous control offline RL benchmarks, with statistically significant improvements over baselines (Wilcoxon signed-rank, p < 0.01). We also compute correlation to the environment rewards as a proxy for measuring alignment to the underlying preference criteria. We show that the SARA computed rewards display higher correlation across noise rates compared to baselines.

📄 PDF Abstract BibTeX arXiv:2506.12529

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data

2025-04-14 · Shuai Zhao, Linchao Zhu, Yi Yang

Large language models~(LLMs) are expected to be helpful, harmless, and honest. In various alignment scenarios, such as general human preference, safety, and confidence alignment, binary preference data collection and rew…

Language ModelingLanguage Modelling

Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation

2026-01-18 · Jinmei Liu, Haoru Li, Zhenhong Sun, Chaofeng Chen 외 arxiv

Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks. A fundame…

Reinforcement LearningImage Generation

A First-Order Logic-Based Alternative to Reward Models in RLHF

2025-12-16 · Chunjin Jian, Xinhua Zhu arxiv

Reinforcement Learning from Human Feedback (RLHF) plays a crucial role in aligning large language models (LLMs) with human values and preferences. However, the quality and stability of the trained reward model largely de…

Reinforcement Learning

Prior Constraints-based Reward Model Training for Aligning Large Language Models

2024-04-01 · Hang Zhou, Chenglong Wang, Yimin Hu, Tong Xiao 외

Reinforcement learning with human feedback for aligning large language models (LLMs) trains a reward model typically using ranking loss with comparison pairs.However, the training procedure suffers from an inherent probl…

reinforcement-learningReinforcement Learning

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

2025-07-02 · Chris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He 외 arxiv

Despite the critical role of reward models (RMs) in Reinforcement Learning from Human Feedback (RLHF), current state-of-the-art open RMs perform poorly on most existing evaluation benchmarks, failing to capture nuanced h…

Reinforcement Learning