paper-with-me

홈 › Papers

Sample-efficient LLM Optimization with Reset Replay

2025-08-08 · Zichuan Liu, Jinyu Wang, Lei Song, Jiang Bian arxiv

Recent advancements in LLM post-training, particularly through reinforcement learning and preference optimization, are key to boosting their reasoning capabilities. However, these methods often suffer from low sample efficiency and a susceptibility to primacy bias, a phenomenon where overfitting to initial experiences diminishes network plasticity and damages the learning process. To address these challenges, we introduce LLM optimization with Reset Replay (LoRR), a general and powerful plugin for enhancing sample efficiency in preference-based optimization. Its core mechanism enables high-replay training to maximize the utility of each data batch. To mitigate overfitting, LoRR orchestrates a periodic reset strategy that reuses the initial data and policy to maintain network plasticity, and further adopts a hybrid optimization objective to better exploit training data. Extensive experiments show that LoRR significantly boosts the performance of various preference optimization methods on both mathematical and general reasoning benchmarks. Notably, an iterative DPO framework augmented with LoRR achieves comparable performance on challenging math tasks, rivaling many complex or computationally expensive baselines. Our findings highlight that LoRR offers a practical and sample-efficient paradigm from limited offline data, unlocking greater performance with minimal changes to existing post-training workflows.

📄 PDF Abstract BibTeX arXiv:2508.06412

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Bilevel Coreset Selection in Continual Learning: A New Formulation and Algorithm

2023-09-21 · NeurIPS 2023 11

Coreset is a small set that provides a data summary for a large dataset, such that training solely on the small set achieves competitive performance compared with a large dataset. In rehearsal-based continual learning, t…

Sample-Efficient and Safe Deep Reinforcement Learning via Reset Deep Ensemble Agents

2023-10-31 · NeurIPS 2023 11

Deep reinforcement learning (RL) has achieved remarkable success in solving complex tasks through its integration with deep neural networks (DNNs) as function approximators. However, the reliance on DNNs has introduced a…

Deep Reinforcement LearningEnsemble LearningReinforcement Learning (RL)

GCR: Gradient Coreset Based Replay Buffer Selection For Continual Learning

2021-11-18 · CVPR 2022 1 · Rishabh Tiwari, KrishnaTeja Killamsetty, Rishabh Iyer, Pradeep Shenoy

Continual learning (CL) aims to develop techniques by which a single model adapts to an increasing number of tasks encountered sequentially, thereby potentially leveraging learnings across tasks in a resource-efficient m…

Continual Learning

Selective experience replay compression using coresets for lifelong deep reinforcement learning in medical imaging

2023-02-22 · Guangyao Zheng, Samson Zhou, Vladimir Braverman, Michael A. Jacobs 외

Selective experience replay is a popular strategy for integrating lifelong learning with deep reinforcement learning. Selective experience replay aims to recount selected experiences from previous tasks to avoid catastro…

Brain Tumor SegmentationDeep Reinforcement LearningLifelong learningTumor Segmentation

PDAC: Efficient Coreset Selection for Continual Learning via Probability Density Awareness

2025-11-12 · Junqi Gao, Zhichang Guo, Dazhi Zhang, Yao Li 외 arxiv

Rehearsal-based Continual Learning (CL) maintains a limited memory buffer to store replay samples for knowledge retention, making these approaches heavily reliant on the quality of the stored samples. Current Rehearsal-b…

Continual Learning