Sailing AI by the Stars: A Survey of Learning from Rewards in Post-Training and Test-Time Scaling of Large Language Models
Recent developments in Large Language Models (LLMs) have shifted from pre-training scaling to post-training and test-time scaling. Across these developments, a key unified paradigm has arisen: Learning from Rewards, where reward signals act as the guiding stars to steer LLM behavior. It has underpinned a wide range of prevalent techniques, such as reinforcement learning (in RLHF, DPO, and GRPO), reward-guided decoding, and post-hoc correction. Crucially, this paradigm enables the transition from passive learning from static data to active learning from dynamic feedback. This endows LLMs with aligned preferences and deep reasoning capabilities. In this survey, we present a comprehensive overview of the paradigm of learning from rewards. We categorize and analyze the strategies under this paradigm across training, inference, and post-inference stages. We further discuss the benchmarks for reward models and the primary applications. Finally we highlight the challenges and future directions. We maintain a paper collection at https://github.com/bobxwu/learning-from-rewards-llm-papers.
Code (1)
Tasks
Active LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Deploying Reinforcement Learning in Water Transport
In this project, we deployed various types of reinforcement learning algorithms to resolves the rewards maximization problem in water sailing by reaching the destination with the highest priority using the smallest numbe…
Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Investigation of stellar magnetic activity using variational autoencoder based on low-resolution spectroscopic survey
We apply the variational autoencoder (VAE) to the LAMOST-K2 low-resolution spectra to detect the magnetic activity of the stars in the K2 field. After the training on the spectra of the selected inactive stars, the VAE m…
TripletAutomatic classification of eclipsing binary stars using deep learning methods
In the last couple of decades, tremendous progress has been achieved in developing robotic telescopes and, as a result, sky surveys (both terrestrial and space) have become the source of a substantial amount of new obser…
ClassificationDeep LearningPhotometric identification of compact galaxies, stars and quasars using multiple neural networks
We present MargNet, a deep learning-based classifier for identifying stars, quasars and compact galaxies using photometric parameters and images from the Sloan Digital Sky Survey (SDSS) Data Release 16 (DR16) catalogue. …
Deep LearningFeature EngineeringSurveyBeaming Binaries - a New Observational Category of Photometric Binary Stars
The new photometric space-borne survey missions CoRoT and Kepler will be able to detect minute flux variations in binary stars due to relativistic beaming caused by the line-of-sight motion of their components. In all bu…