paper-with-me

홈 › Papers

Sailing AI by the Stars: A Survey of Learning from Rewards in Post-Training and Test-Time Scaling of Large Language Models

2025-05-05 · Xiaobao Wu

Recent developments in Large Language Models (LLMs) have shifted from pre-training scaling to post-training and test-time scaling. Across these developments, a key unified paradigm has arisen: Learning from Rewards, where reward signals act as the guiding stars to steer LLM behavior. It has underpinned a wide range of prevalent techniques, such as reinforcement learning (in RLHF, DPO, and GRPO), reward-guided decoding, and post-hoc correction. Crucially, this paradigm enables the transition from passive learning from static data to active learning from dynamic feedback. This endows LLMs with aligned preferences and deep reasoning capabilities. In this survey, we present a comprehensive overview of the paradigm of learning from rewards. We categorize and analyze the strategies under this paradigm across training, inference, and post-inference stages. We further discuss the benchmarks for reward models and the primary applications. Finally we highlight the challenges and future directions. We maintain a paper collection at https://github.com/bobxwu/learning-from-rewards-llm-papers.

📄 PDF Abstract BibTeX arXiv:2505.02686

Code (1)

bobxwu/learning-from-rewards-llm-papers 공식 구현

Tasks

Active Learning

Methods 이 논문이 사용한 방법론

DPO 설명 없음

Similar Papers 제목 키워드 기반

Deploying Reinforcement Learning in Water Transport

2020-12-14 · CUHK Course IERG5350 2020 12 · Donald P. H. Wong, Tsz Kui Chow

In this project, we deployed various types of reinforcement learning algorithms to resolves the rewards maximization problem in water sailing by reaching the destination with the highest priority using the smallest numbe…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Investigation of stellar magnetic activity using variational autoencoder based on low-resolution spectroscopic survey

2022-06-15 · Yue Xiang, Shenghong Gu, Dongtao Cao

We apply the variational autoencoder (VAE) to the LAMOST-K2 low-resolution spectra to detect the magnetic activity of the stars in the K2 field. After the training on the spectra of the selected inactive stars, the VAE m…

Triplet

Automatic classification of eclipsing binary stars using deep learning methods

2021-08-03 · Michal Čokina, Viera Maslej-Krešňáková, Peter Butka, Štefan Parimucha

In the last couple of decades, tremendous progress has been achieved in developing robotic telescopes and, as a result, sky surveys (both terrestrial and space) have become the source of a substantial amount of new obser…

ClassificationDeep Learning

Photometric identification of compact galaxies, stars and quasars using multiple neural networks

2022-11-15 · Siddharth Chaini, Atharva Bagul, Anish Deshpande, Rishi Gondkar 외

We present MargNet, a deep learning-based classifier for identifying stars, quasars and compact galaxies using photometric parameters and images from the Sloan Digital Sky Survey (SDSS) Data Release 16 (DR16) catalogue. …

Deep LearningFeature EngineeringSurvey

Beaming Binaries - a New Observational Category of Photometric Binary Stars

2007-08-15 · Shay Zucker, Tsevi Mazeh, Tal Alexander

The new photometric space-borne survey missions CoRoT and Kepler will be able to detect minute flux variations in binary stars due to relativistic beaming caused by the line-of-sight motion of their components. In all bu…