paper-with-me

홈 › Papers

Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting

2025-11-17 · Yunhun Nam, Jaehyung Kim, Jongheon Jeong arxiv

Language models (LMs) are often adapted through supervised fine-tuning (SFT) to specialize their capabilities for downstream tasks. However, in typical scenarios where the fine-tuning data is limited, e.g., compared to pre-training, SFT can lead LMs to overfit, causing them to rely on spurious patterns within the target task or to compromise other broadly useful capabilities as a side effect of narrow specialization. In this paper, we propose Learning-from-the-Undesirable (LfU), a simple yet effective regularization scheme for SFT to mitigate overfitting issues when fine-tuning LMs with limited data. Specifically, we aim to regularize the fine-tuning process to favor solutions that are resilient to "undesirable" model updates, e.g., gradient ascent steps that steer the model toward undesirable behaviors. To this end, we propose a novel form of consistency regularization that directly aligns internal representations of the model with those after an undesirable update. By leveraging representation-level data augmentation through undesirable updates, LfU effectively promotes generalization under limited data. Our experiments on diverse LM downstream tasks show that LfU serves as an effective prior that enhances adaptability while preserving pretrained knowledge. For example, our LM from LfU achieves a 16.8% average improvement on math tasks compared to vanilla SFT on the same dataset, where the latter even leads to degraded performance on those tasks. Furthermore, LfU exhibits improved robustness to prompt variations, e.g., yielding a 92.1% lower standard deviation in output performances compared to SFT, highlighting its versatile effects.

📄 PDF Abstract BibTeX arXiv:2511.13052

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

Fortuitous Forgetting in Connectionist Networks

2022-02-01 · ICLR 2022 4 · Hattie Zhou, Ankit Vani, Hugo Larochelle, Aaron Courville

Forgetting is often seen as an unwanted characteristic in both human and machine learning. However, we propose that forgetting can in fact be favorable to learning. We introduce "forget-and-relearn" as a powerful paradig…

image-classificationImage Classification

SAFT: Towards Out-of-Distribution Generalization in Fine-Tuning

2024-07-03 · Bac Nguyen, Stefan Uhlich, Fabien Cardinaux, Lukas Mauch 외

Handling distribution shifts from training data, known as out-of-distribution (OOD) generalization, poses a significant challenge in the field of machine learning. While a pre-trained vision-language model like CLIP has …

Few-Shot LearningGeneral KnowledgeLanguage ModelingLanguage Modelling+1

DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher

2026-01-29 · Yisheng Zhong, Zhengbang Yang, Zhuangdi Zhu arxiv

LLM unlearning is a technique to remove the impacts of undesirable knowledge from the model without retraining from scratch, which is indispensable towards trustworthy AI. Existing unlearning methods face significant lim…

VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models

2026-04-05 · Ravi Ranjan, Agoritsa Polyzou arxiv

Vision-language-action (VLA) models are emerging as embodied foundation models for robotic manipulation, but their deployment introduces a new unlearning challenge: removing unsafe, spurious, or privacy-sensitive behavio…

Bayesian Parameter-Efficient Fine-Tuning for Overcoming Catastrophic Forgetting

2024-02-19 · Haolin Chen, Philip N. Garner

We are motivated primarily by the adaptation of text-to-speech synthesis models; however we argue that more generic parameter-efficient fine-tuning (PEFT) is an appropriate framework to do such adaptation. Nevertheless, …

Language ModelingLanguage Modellingparameter-efficient fine-tuningSpeech Synthesis+3