Reinforcement Pre-Training
In this work, we introduce Reinforcement Pre-Training (RPT) as a new scaling paradigm for large language models and reinforcement learning (RL). Specifically, we reframe next-token prediction as a reasoning task trained using RL, where it receives verifiable rewards for correctly predicting the next token for a given context. RPT offers a scalable method to leverage vast amounts of text data for general-purpose RL, rather than relying on domain-specific annotated answers. By incentivizing the capability of next-token reasoning, RPT significantly improves the language modeling accuracy of predicting the next tokens. Moreover, RPT provides a strong pre-trained foundation for further reinforcement fine-tuning. The scaling curves show that increased training compute consistently improves the next-token prediction accuracy. The results position RPT as an effective and promising scaling paradigm to advance language model pre-training.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Non-Robust Feature Mapping in Deep Reinforcement Learning
Adversarial perturbations to state observations can dramatically degrade the performance of deep reinforcement learning policies, and thus raise concerns regarding the robustness of deep reinforcement learning agents. A …
Atari GamesDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1Reinforcement Networks: novel framework for collaborative Multi-Agent Reinforcement Learning tasks
Modern AI systems often comprise multiple learnable components that can be naturally organized as graphs. A central challenge is the end-to-end training of such systems without restrictive architectural or training assum…
Multi-agent Reinforcement LearningPseudo-Model-Free Hedging for Variable Annuities via Deep Reinforcement Learning
This paper proposes a two-phase deep reinforcement learning approach, for hedging variable annuity contracts with both GMMB and GMDB riders, which can address model miscalibration in Black-Scholes financial and constant …
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Benchmarking Robustness of Deep Reinforcement Learning approaches to Online Portfolio Management
Deep Reinforcement Learning approaches to Online Portfolio Selection have grown in popularity in recent years. The sensitive nature of training Reinforcement Learning agents implies a need for extensive efforts in market…
BenchmarkingDeep Reinforcement LearningManagementreinforcement-learning+1QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning
Deep reinforcement learning continues to show tremendous potential in achieving task-level autonomy, however, its computational and energy demands remain prohibitively high. In this paper, we tackle this problem by apply…
Decision MakingDeep Reinforcement LearningQuantizationreinforcement-learning+2