paper-with-me

홈 › Papers

A Survey on Progress in LLM Alignment from the Perspective of Reward Design

2025-05-05 · Miaomiao Ji, Yanqiu Wu, Zhibin Wu, Shoujin Wang, Jian Yang, Mark Dras, Usman Naseem

The alignment of large language models (LLMs) with human values and intentions represents a core challenge in current AI research, where reward mechanism design has become a critical factor in shaping model behavior. This study conducts a comprehensive investigation of reward mechanisms in LLM alignment through a systematic theoretical framework, categorizing their development into three key phases: (1) feedback (diagnosis), (2) reward design (prescription), and (3) optimization (treatment). Through a four-dimensional analysis encompassing construction basis, format, expression, and granularity, this research establishes a systematic classification framework that reveals evolutionary trends in reward modeling. The field of LLM alignment faces several persistent challenges, while recent advances in reward design are driving significant paradigm shifts. Notable developments include the transition from reinforcement learning-based frameworks to novel optimization paradigms, as well as enhanced capabilities to address complex alignment scenarios involving multimodal integration and concurrent task coordination. Finally, this survey outlines promising future research directions for LLM alignment through innovative reward design strategies.

📄 PDF Abstract BibTeX arXiv:2505.02666

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

2026-07-22 · Jianshu Zhang, Keliang Wu, Haoran Lu, Anbang Liu 외 hf

Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making pr…

AI Alignment through a Game-theoretic Lens: A Survey

2026-08-28 · Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee 외 arxiv

As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in impro…

A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

2025-10-09 · Congmin Zheng, Jiachen Zhu, Zhuoying Ou, Yuxiang Chen 외 arxiv

Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward models (ORMs) that judge only final answers. Process Reward Models(PRMs) addres…

Reinforcement LearningMultimodal Reasoning

SurveyLM: A platform to explore emerging value perspectives in augmented language models' behaviors

2023-08-01 · Steve J. Bickley, Ho Fai Chan, Bang Dao, Benno Torgler 외

This white paper presents our work on SurveyLM, a platform for analyzing augmented language models' (ALMs) emergent alignment behaviors through their dynamically evolving attitude and value perspectives in complex social…

Survey

Reinforcement Learning from Human Feedback: A Statistical Perspective

2026-04-02 · Pangpang Liu, Chengchun Shi, Will Wei Sun arxiv

Reinforcement learning from human feedback (RLHF) has emerged as a central framework for aligning large language models (LLMs) with human preferences. Despite its practical success, RLHF raises fundamental statistical qu…

Reinforcement LearningActive Learning