paper-with-me

Papers

Governance Challenges in Reinforcement Learning from Human Feedback: Evaluator Rationality and Reinforcement Stability

2025-04-17 · Dana Alsagheer, Abdulrahman Kamal, Mohammad Kamal, Weidong Shi

Reinforcement Learning from Human Feedback (RLHF) is central in aligning large language models (LLMs) with human values and expectations. However, the process remains susceptible to governance challenges, including evaluator bias, inconsistency, and the unreliability of feedback. This study examines how the cognitive capacity of evaluators, specifically their level of rationality, affects the stability of reinforcement signals. A controlled experiment comparing high-rationality and low-rationality participants reveals that evaluators with higher rationality scores produce significantly more consistent and expert-aligned feedback. In contrast, lower-rationality participants demonstrate considerable variability in their reinforcement decisions ($p < 0.01$). To address these challenges and improve RLHF governance, we recommend implementing evaluator pre-screening, systematic auditing of feedback consistency, and reliability-weighted reinforcement aggregation. These measures enhance the fairness, transparency, and robustness of AI alignment pipelines.

📄 PDF Abstract BibTeX arXiv:2504.13972

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Similar Papers 제목 키워드 기반

Human Implicit Preference-Based Policy Fine-tuning for Multi-Agent Reinforcement Learning in USV Swarm

2025-03-05 · HyeonJun Kim, Kanghoon Lee, Junho Park, Jiachen Li 외

Multi-Agent Reinforcement Learning (MARL) has shown promise in solving complex problems involving cooperation and competition among agents, such as an Unmanned Surface Vehicle (USV) swarm used in search and rescue, surve…

Collision AvoidanceFairnessLanguage ModelingLanguage Modelling+4

When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback

2024-02-27 · Leon Lang, Davis Foote, Stuart Russell, Anca Dragan 외

Past analyses of reinforcement learning from human feedback (RLHF) assume that the human evaluators fully observe the environment. What happens when human feedback is based only on partial observations? We formally defin…

HRLAIF: Improvements in Helpfulness and Harmlessness in Open-domain Reinforcement Learning From AI Feedback

2024-03-13 · Ang Li, Qiugen Xiao, Peng Cao, Jian Tang 외

Reinforcement Learning from AI Feedback (RLAIF) has the advantages of shorter annotation cycles and lower costs over Reinforcement Learning from Human Feedback (RLHF), making it highly efficient during the rapid strategy…

Language ModellingLarge Language ModelRed Teamingreinforcement-learning+2

Reinforcement Learning from User Feedback

2025-05-20 · Eric Han, Jun Chen, Karthik Abinav Sankararaman, Xiaoliang Peng 외

As large language models (LLMs) are increasingly deployed in diverse user facing applications, aligning them with real user preferences becomes essential. Existing methods like Reinforcement Learning from Human Feedback …

reinforcement-learningReinforcement Learning

A Large Language Model-Driven Reward Design Framework via Dynamic Feedback for Reinforcement Learning

2024-10-18 · Shengjie Sun, Runze Liu, Jiafei Lyu, Jing-Wen Yang 외

Large Language Models (LLMs) have shown significant potential in designing reward functions for Reinforcement Learning (RL) tasks. However, obtaining high-quality reward code often involves human intervention, numerous L…

Language ModelingLanguage ModellingLarge Language ModelReinforcement Learning (RL)