paper-with-me

홈 › Papers

Effective Skill Unlearning through Intervention and Abstention

2025-03-27 · Yongce Li, Chung-En Sun, Tsui-Wei Weng

Large language Models (LLMs) have demonstrated remarkable skills across various domains. Understanding the mechanisms behind their abilities and implementing controls over them is becoming increasingly important for developing better models. In this paper, we focus on skill unlearning in LLMs, specifically unlearning a particular skill while retaining their overall capabilities. We introduce two lightweight, training-free machine skill unlearning techniques for LLMs. First, we observe that the pre-activation distribution of neurons in each Feed-Forward Layer (FFL) differs when the model demonstrates different skills. Additionally, we find that queries triggering the same skill cluster within the FFL key space and can be separated from other queries using a hypercube. Based on these observations, we propose two lightweight, training-free skill unlearning methods via \textit{intervention} and \textit{abstention} respectively: \texttt{Neuron Adjust} and \texttt{Key Space Detection}. We evaluate our methods on unlearning math-solving, Python-coding, and comprehension skills across seven different languages. The results demonstrate their strong unlearning capabilities for the designated skills. Specifically, \texttt{Key Space Detection} achieves over 80\% relative performance drop on the forgetting skill and less than 10\% relative performance drop on other skills and the model's general knowledge (MMLU) for most unlearning tasks. Our code is available at https://github.com/Trustworthy-ML-Lab/effective_skill_unlearning

📄 PDF Abstract BibTeX arXiv:2503.21730

Code (1)

trustworthy-ml-lab/effective_skill_unlearning 공식 구현 pytorch

Tasks

General KnowledgeMathMMLU

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Why Instruction-Based Unlearning Fails in Diffusion Models?

2026-04-02 · Zeliang Zhang, Rui Sun, Jiani Liu, Qi Wu 외 arxiv

Instruction-based unlearning has proven effective for modifying the behavior of large language models at inference time, but whether this paradigm extends to other generative models remains unclear. In this work, we inve…

Image Generation

Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning

2026-05-26 · Junkai Chen, Yuhao He, Junxiang You, Ruiqi Liu 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable progress on vision-language tasks, but they may also memorize and expose sensitive or restricted knowledge, raising concerns about privacy and broader saf…

To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models

2025-10-15 · Anna Hedström, Salim I. Amoukou, Tom Bewley, Saumitra Mishra 외 arxiv

We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely o…

Auditing Language Model Unlearning via Information Decomposition

2026-01-21 · Anmol Goel, Alan Ritter, Iryna Gurevych arxiv

We expose a critical limitation in current approaches to machine unlearning in language models: despite the apparent success of unlearning algorithms, information about the forgotten data remains linearly decodable from …

CLUE: Conflict-guided Localization for LLM Unlearning Framework

2025-09-25 · Hang Chen, Jiaying Zhu, Xinyu Yang, Wenya Wang arxiv

The LLM unlearning aims to eliminate the influence of undesirable data without affecting causally unrelated information. This process typically involves using a forget set to remove target information, alongside a retain…