paper-with-me

Papers

VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback

2024-09-27 · Guoxi Zhang, Jiuding Duan

This paper addresses the cost-efficiency aspect of Reinforcement Learning from Human Feedback (RLHF). RLHF leverages datasets of human preferences over outputs of large language models (LLM)s to instill human expectations into LLMs. Although preference annotation comes with a monetized cost, the economic utility of a preference dataset has not been considered by far. What exacerbates this situation is that, given complex intransitive or cyclic relationships in preference datasets, existing algorithms for fine-tuning LLMs are still far from capturing comprehensive preferences. This raises severe cost-efficiency concerns in production environments, where preference data accumulate over time. In this paper, we discuss the fine-tuning of LLMs as a monetized economy and introduce an auction mechanism to improve the efficiency of preference data collection in dollar terms. We show that introducing an auction mechanism can play an essential role in enhancing the cost-efficiency of RLHF, while maintaining satisfactory model performance. Experimental results demonstrate that our proposed auction-based protocol is cost-effective for fine-tuning LLMs concentrating on high-quality feedback.

📄 PDF Abstract BibTeX arXiv:2409.18417

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning to Construct Knowledge through Sparse Reference Selection with Reinforcement Learning

2025-09-07 · Shao-An Yin arxiv

The rapid expansion of scientific literature makes it increasingly difficult to acquire new knowledge, particularly in specialized domains where reasoning is complex, full-text access is restricted, and target references…

Reinforcement Learning

Toward Global Intent Inference for Human Motion by Inverse Reinforcement Learning

2026-03-08 · Sarmad Mehrdad, Maxime Sabbah, Vincent Bonnet, Ludovic Righetti arxiv

This paper investigates whether a single, unified cost function can explain and predict human reaching movements, in contrast with existing approaches that rely on subject- or posture-specific optimization criteria. Usin…

Reinforcement Learning

Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning

2024-07-02 · Yifang Chen, Shuohang Wang, ZiYi Yang, Hiteshi Sharma 외

Reinforcement learning with human feedback (RLHF), as a widely adopted approach in current large language model pipelines, is \textit{bottlenecked by the size of human preference data}. While traditional methods rely on …

Active LearningLanguage ModellingLarge Language ModelMMLU

Towards Human-Centered Construction Robotics: A Reinforcement Learning-Driven Companion Robot for Contextually Assisting Carpentry Workers

2024-03-27 · Yuning Wu, Jiaying Wei, Jean Oh, Daniel Cardoso Llach

In the dynamic construction industry, traditional robotic integration has primarily focused on automating specific tasks, often overlooking the complexity and variability of human aspects in construction workflows. This …

Reinforcement Learning (RL)

MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction

2026-02-17 · Qiang Zhang, Jiahao Ma, Peiran Liu, Shuai Shi 외 arxiv

Humanoid motion control has witnessed significant breakthroughs in recent years, with deep reinforcement learning (RL) emerging as a primary catalyst for achieving complex, human-like behaviors. However, the high dimensi…

Reinforcement LearningMotion Synthesis