paper-with-me

홈 › Papers

Best Practice Critic Optimization

2026-08-24 · Penghui Qi, Xiangxin Zhou, Wee Sun Lee arxiv

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes are often unstable. We study this instability and develop Best Practice Critic Optimization (BPCO), a recipe that combines DPPO, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive generalized advantage estimation. Because the critic is used only during training, BPCO can also condition it on reward-defining information, such as a reference answer or grading rubric, that is hidden from the policy. Controlled experiments isolate the effect of each design choice. Across mathematical reasoning tasks with models ranging from 1.5B parameters to 30B-A3B mixtures of experts, BPCO improves a strong critic-based baseline consistently, and matches or exceeds a group-based baseline while sampling one response per prompt. The same recipe also improves learning with rubric-based rewards. These results show that a carefully designed critic provides a reliable alternative to group-relative advantage estimation. Code is available at https://github.com/QPHutu/golden_critic.

📄 PDF Abstract BibTeX arXiv:2608.23566

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

Best Practices for Machine Learning Systems: An Industrial Framework for Analysis and Optimization

2023-06-09 · Georgios Christos Chouliaras, Kornel Kiełczewski, Amit Beka, David Konopnicki 외

In the last few years, the Machine Learning (ML) and Artificial Intelligence community has developed an increasing interest in Software Engineering (SE) for ML Systems leading to a proliferation of best practices, rules,…

High-probability Bounds for Non-Convex Stochastic Optimization with Heavy Tails

2021-06-28 · NeurIPS 2021 12 · Ashok Cutkosky, Harsh Mehta

We consider non-convex stochastic optimization using first-order algorithms for which the gradient estimates may have heavy tails. We show that a combination of gradient clipping, momentum, and normalized gradient descen…

Stochastic OptimizationVocal Bursts Intensity Prediction

Risk-averse Stochastic Optimization for Farm Management Practices and Cultivar Selection Under Uncertainty

2022-07-17 · Faezeh Akhavizadegan, Javad Ansarifar, Lizhi Wang, Sotirios V. Archontoulis

Optimizing management practices and selecting the best cultivar for planting play a significant role in increasing agricultural food production and decreasing environmental footprint. In this study, we develop optimizati…

Bayesian OptimizationManagementStochastic Optimization

Best Practices for Text Annotation with Large Language Models

2024-02-05 · Petter Törnberg

Large Language Models (LLMs) have ushered in a new era of text annotation, as their ease-of-use, high accuracy, and relatively low costs have meant that their use has exploded in recent months. However, the rapid growth …

Model SelectionPrompt Engineeringtext annotation

Benchmarking in Optimization: Best Practice and Open Issues

2020-07-07 · Thomas Bartz-Beielstein, Carola Doerr, Daan van den Berg, Jakob Bossek 외

This survey compiles ideas and recommendations from more than a dozen researchers with different backgrounds and from different institutes around the world. Promoting best practice in benchmarking is its main goal. The a…

Benchmarking