paper-with-me

홈 › Papers

Reward Modeling with Weak Supervision for Language Models

2024-10-28 · Ben Hauptvogel, Malte Ostendorff, Georg Rehm, Sebastian Möller

Recent advancements in large language models (LLMs) have led to their increased application across various tasks, with reinforcement learning from human feedback (RLHF) being a crucial part of their training to align responses with user intentions. In the RLHF process, a reward model is trained using responses preferences determined by human labelers or AI systems, which then refines the LLM through reinforcement learning. This work introduces weak supervision as a strategy to extend RLHF datasets and enhance reward model performance. Weak supervision employs noisy or imprecise data labeling, reducing reliance on expensive manually labeled data. By analyzing RLHF datasets to identify heuristics that correlate with response preference, we wrote simple labeling functions and then calibrated a label model to weakly annotate unlabeled data. Our evaluation show that while weak supervision significantly benefits smaller datasets by improving reward model performance, its effectiveness decreases with larger, originally labeled datasets. Additionally, using an LLM to generate and then weakly label responses offers a promising method for extending preference data.

📄 PDF Abstract BibTeX arXiv:2410.20869

Code (1)

DFKI-NLP/weak-supervision-rlhf 공식 구현

Tasks

reinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning

2026-03-09 · Chi-Min Chan, Ehsan Hajiramezanali, Xiner Li, Edward De Brouwer 외 arxiv

In scientific reasoning tasks, the veracity of the reasoning process is as critical as the final outcome. While Process Reward Models (PRMs) offer a solution to the coarse-grained supervision problems inherent in Outcome…

When Can LLMs Learn to Reason with Weak Supervision?

2026-04-20 · Salman Rahman, Jingyan Shen, Anna Mordvina, Hamid Palangi 외 arxiv

Large language models have achieved significant reasoning improvements through reinforcement learning with verifiable rewards (RLVR). Yet as model capabilities grow, constructing high-quality reward signals becomes incre…

Reinforcement Learning

From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models

2025-03-08 · Muzhi Dai, Jiashuo Sun, Zhiyuan Zhao, Shixuan Liu 외

Aligning large vision-language models (LVLMs) with human preferences is challenging due to the scarcity of fine-grained, high-quality, and multimodal preference data without human annotations. Existing methods relying on…

Image CaptioningLanguage ModelingLanguage ModellingLarge Language Model

Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

2023-12-14 · Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 외

Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior - for example, to evaluate whether a model faithfully followed instru…

Policy Learning Using Weak Supervision

2020-10-05 · NeurIPS 2021 12 · Jingkang Wang, Hongyi Guo, Zhaowei Zhu, Yang Liu

Most existing policy learning solutions require the learning agents to receive high-quality supervision signals such as well-designed rewards in reinforcement learning (RL) or high-quality expert demonstrations in behavi…

Reinforcement Learning (RL)