paper-with-me

홈 › Papers

Reinforcement Learning from Meta-Evaluation: Aligning Language Models Without Ground-Truth Labels

2026-01-29 · Micah Rentschler, Jesse Roberts arxiv

Most reinforcement learning (RL) methods for training large language models (LLMs) require ground-truth labels or task-specific verifiers, limiting scalability when correctness is ambiguous or expensive to obtain. We introduce Reinforcement Learning from Meta-Evaluation (RLME), which optimizes a generator using reward derived from an evaluator's answers to natural-language meta-questions (e.g., "Is the answer correct?" or "Is the reasoning logically consistent?"). RLME treats the evaluator's probability of a positive judgment as a reward and updates the generator via group-relative policy optimization, enabling learning without labels. Across a suite of experiments, we show that RLME achieves accuracy and sample efficiency comparable to label-based training, enables controllable trade-offs among multiple objectives, steers models toward reliable reasoning patterns rather than post-hoc rationalization, and generalizes to open-domain settings where ground-truth labels are unavailable, broadening the domains in which LLMs may be trained with RL.

📄 PDF Abstract BibTeX arXiv:2601.21268

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement Learning

2026-02-12 · Haoran Dang, Cuiling Lan, Hai Wan, Xibin Zhao 외 arxiv

Temperature is a crucial hyperparameter in large language models (LLMs), controlling the trade-off between exploration and exploitation during text generation. High temperatures encourage diverse but noisy outputs, while…

Reinforcement LearningMathematical ReasoningText Generation

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

2025-04-23 · Yuran Li, Jama Hussein Mohamud, Chongren Sun, Di wu 외

Large language models (LLMs) are being widely applied across various fields, but as tasks become more complex, evaluating their responses is increasingly challenging. Compared to human evaluators, the use of LLMs to supp…

MetaAlign: Align Large Language Models with Diverse Preferences during Inference Time

2024-10-18 · Mozhi Zhang, Pengyu Wang, Chenkun Tan, Mianqiu Huang 외

Large Language Models (LLMs) acquire extensive knowledge and remarkable abilities from extensive text corpora, making them powerful tools for various applications. To make LLMs more usable, aligning them with human prefe…

Diversity

Revisiting Meta-evaluation for Grammatical Error Correction

2024-03-05 · Masamune Kobayashi, Masato Mita, Mamoru Komachi

Metrics are the foundation for automatic evaluation in grammatical error correction (GEC), with their evaluation of the metrics (meta-evaluation) relying on their correlation with human judgments. However, conventional m…

Grammatical Error CorrectionSentence

Privately Aligning Language Models with Reinforcement Learning

2023-10-25 · Fan Wu, Huseyin A. Inan, Arturs Backurs, Varun Chandrasekaran 외

Positioned between pre-training and user deployment, aligning large language models (LLMs) through reinforcement learning (RL) has emerged as a prevailing strategy for training instruction following-models such as ChatGP…

Instruction FollowingPrivacy Preservingreinforcement-learningReinforcement Learning+2