paper-with-me

홈 › Papers

Mechanism-Guided Selective Unlearning for RLVR-Induced Reasoning

2026-06-17 · Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou arxiv

We propose MAST (Mechanism-Aligned Selective Targeting), a mechanism-guided method for unlearning RLVR-induced reasoning with substantially lower collateral damage than standard full-parameter updates. In matched SFT/RLVR checkpoints on Qwen2.5-Math-1.5B and Qwen3-1.7B-Base, the SFT-to-RLVR increment differs sharply from the SFT update in token-level delta-log-probability, and full-parameter gradient ascent forgets only by damaging retain MATH and GSM8K. MAST ranks attention-projection tensors by off-principal energy, update magnitude, and forget-gradient coupling magnitude, then updates only the top-ranked subset. On the primary model, MAST induces statistically significant target forgetting (MATH forget 45/150 to 37/150; McNemar p=0.0078) while preserving GSM8K (+0.8 pp) and MATH retain (-0.5 pp). The advantage reproduces across seeds, NPO/SimNPO objectives, and Qwen3, where MAST preserves GSM8K while full-parameter unlearning collapses it.

📄 PDF Abstract BibTeX arXiv:2606.19222

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models

2025-12-19 · Wei Qian, Chenxu Zhao, Yangyi Li, Mengdi Huai arxiv

The rapid advancements in artificial intelligence (AI) have primarily focused on the process of learning from data to acquire knowledgeable learning systems. As these systems are increasingly deployed in critical areas, …

Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

2026-07-30 · Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj, Gjergji Kasneci arxiv

Machine unlearning seeks to selectively remove specific knowledge from trained language models without full retraining, a growing necessity under privacy regulations such as GDPR and the EU AI Act. Recent work has reform…

Reinforcement Learning

IMU: Influence-guided Machine Unlearning

2025-08-03 · Xindi Fan, Jing Wu, Mingyi Zhou, Pengwei Liang 외 arxiv

Machine Unlearning (MU) aims to selectively erase the influence of specific data points from pretrained models. However, most existing MU methods rely on the retain set to preserve model utility, which is often impractic…

Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds

2026-08-27 · Waqas Khan, Tabinda Sarwar, Jingyue Cong, Xun Yi 외 arxiv

Enterprises fine-tune language models on proprietary data that may later require removal due to privacy, contractual, or compliance obligations. Selective unlearning removes requested knowledge while preserving model uti…

Selective Fine-Tuning for Targeted and Robust Concept Unlearning

2026-02-08 · Mansi, Avinash Kori, Francesca Toni, Soteris Demetriou arxiv

Text guided diffusion models are used by millions of users, but can be easily exploited to produce harmful content. Concept unlearning methods aim at reducing the models' likelihood of generating harmful content. Traditi…