paper-with-me

홈 › Papers

WorldPM: Scaling Human Preference Modeling

2025-05-15 · Binghai Wang, Runji Lin, Keming Lu, Le Yu, Zhenru Zhang, Fei Huang, Chujie Zheng, Kai Dang, Yang Fan, Xingzhang Ren, An Yang, Binyuan Hui, Dayiheng Liu, Tao Gui, Qi Zhang, Xuanjing Huang, Yu-Gang Jiang, Bowen Yu, Jingren Zhou, Junyang Lin

Motivated by scaling laws in language modeling that demonstrate how test loss scales as a power law with model and dataset sizes, we find that similar laws exist in preference modeling. We propose World Preference Modeling$ (WorldPM) to emphasize this scaling potential, where World Preference embodies a unified representation of human preferences. In this paper, we collect preference data from public forums covering diverse user communities, and conduct extensive training using 15M-scale data across models ranging from 1.5B to 72B parameters. We observe distinct patterns across different evaluation metrics: (1) Adversarial metrics (ability to identify deceptive features) consistently scale up with increased training data and base model size; (2) Objective metrics (objective knowledge with well-defined answers) show emergent behavior in larger language models, highlighting WorldPM's scalability potential; (3) Subjective metrics (subjective preferences from a limited number of humans or AI) do not demonstrate scaling trends. Further experiments validate the effectiveness of WorldPM as a foundation for preference fine-tuning. Through evaluations on 7 benchmarks with 20 subtasks, we find that WorldPM broadly improves the generalization performance across human preference datasets of varying sizes (7K, 100K and 800K samples), with performance gains exceeding 5% on many key subtasks. Integrating WorldPM into our internal RLHF pipeline, we observe significant improvements on both in-house and public evaluation sets, with notable gains of 4% to 8% in our in-house evaluations.

📄 PDF Abstract BibTeX arXiv:2505.10527

Code (1)

qwenlm/worldpm 공식 구현

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback

2026-02-13 · Amirhossein Afsharrad, Ruida Zhou, Luca Viano, Sanjay Lall 외 arxiv

Reward modeling is crucial for aligning large language models with human preferences, yet current approaches lack a principled mathematical framework for leveraging ordinal preference data. When human annotators provide …

StoryAlign: Evaluating and Training Reward Models for Story Generation

2026-05-06 · Haotian Xia, Hao Peng, Yunjia Qi, Xiaozhi Wang 외 arxiv

Story generation aims to automatically produce coherent, structured, and engaging narratives. Although large language models (LLMs) have significantly advanced text generation, stories generated by LLMs still diverge fro…

Story GenerationText Generation

Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

2025-02-26 · Hao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao 외

Reward models (RMs) are crucial for the training and inference-time scaling up of large language models (LLMs). However, existing reward models primarily focus on human preferences, neglecting verifiable correctness sign…

Instruction Following

Adaptive Preference Scaling for Reinforcement Learning with Human Feedback

2024-06-04 · Ilgee Hong, Zichong Li, Alexander Bukharin, Yixiao Li 외

Reinforcement learning from human feedback (RLHF) is a prevalent approach to align AI systems with human values by learning rewards from human preference data. Due to various reasons, however, such data typically takes t…

reinforcement-learningReinforcement LearningText Generation

A General Language Assistant as a Laboratory for Alignment

2021-12-01 · Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain 외

Given the broad capabilities of large language models, it should be possible to work towards a general-purpose, text-based assistant that is aligned with human values, meaning that it is helpful, honest, and harmless. As…

Imitation Learning