paper-with-me

홈 › Papers

Robust Policy Optimization to Prevent Catastrophic Forgetting

2026-02-09 · Mahdi Sabbaghi, George Pappas, Adel Javanmard, Hamed Hassani arxiv

Large language models are commonly trained through multi-stage post-training: first via RLHF, then fine-tuned for other downstream objectives. Yet even small downstream updates can compromise earlier learned behaviors (e.g., safety), exposing a brittleness known as catastrophic forgetting. This suggests standard RLHF objectives do not guarantee robustness to future adaptation. To address it, most prior work designs downstream-time methods to preserve previously learned behaviors. We argue that preventing this requires pre-finetuning robustness: the base policy should avoid brittle high-reward solutions whose reward drops sharply under standard fine-tuning. We propose Fine-tuning Robust Policy Optimization (FRPO), a robust RLHF framework that optimizes reward not only at the current policy, but across a KL-bounded neighborhood of policies reachable by downstream adaptation. The key idea is to ensure reward stability under policy shifts via a max-min formulation. By modifying GRPO, we develop an algorithm with no extra computation, and empirically show it substantially reduces safety degradation across multiple base models and downstream fine-tuning regimes (SFT and RL) while preserving downstream task performance. We further study a math-focused RL setting, demonstrating that FRPO preserves accuracy under subsequent fine-tuning.

📄 PDF Abstract BibTeX arXiv:2602.08813

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Updating Only Encoders Prevents Catastrophic Forgetting of End-to-End ASR Models

2022-07-01 · Yuki Takashima, Shota Horiguchi, Shinji Watanabe, Paola García 외

In this paper, we present an incremental domain adaptation technique to prevent catastrophic forgetting for an end-to-end automatic speech recognition (ASR) model. Conventional approaches require extra parameters of the …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain Adaptationspeech-recognition+1

Regularizing Trajectories to Mitigate Catastrophic Forgetting

2019-09-25 · Paul Michel, Elisabeth Salesky, Graham Neubig

Regularization-based continual learning approaches generally prevent catastrophic forgetting by augmenting the training loss with an auxiliary objective. However in most practical optimization scenarios with noisy data a…

Continual Learning

On Catastrophic Forgetting and Mode Collapse in Generative Adversarial Networks

2018-07-11 · Hoang Thanh-Tung, Truyen Tran

In this paper, we show that Generative Adversarial Networks (GANs) suffer from catastrophic forgetting even when they are trained to approximate a single target distribution. We show that GAN training is a continual lear…

Continual Learning

Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning

2026-05-10 · Meng Lou, Hanzhong Guo, Linwei Chen, Yizhou Yu arxiv

Recent studies suggest that Reinforcement Fine-Tuning (RFT) is inherently more resilient to catastrophic forgetting than Supervised Fine-Tuning (SFT). However, whether RFT (e.g., GRPO) can effectively overcome forgetting…

class-incremental learningContinual Learning

Continuous learning of spiking networks trained with local rules

2021-11-18 · Dmitry Antonov, Kirill Sviatov, Sergey Sukhov

Artificial neural networks (ANNs) experience catastrophic forgetting (CF) during sequential learning. In contrast, the brain can learn continuously without any signs of catastrophic forgetting. Spiking neural networks (S…