paper-with-me

홈 › Papers

WARP: Weight Teleportation for Attack-Resilient Unlearning Protocols

2025-11-29 · Mohammad M Maheri, Xavier Cadet, Peter Chin, Hamed Haddadi arxiv

Approximate machine unlearning aims to efficiently remove the influence of specific data points from a trained model, offering a practical alternative to full retraining. However, it introduces privacy risks: an adversary with access to pre- and post-unlearning models can exploit their differences for membership inference or data reconstruction. We show these vulnerabilities arise from two factors: large gradient norms of forget-set samples and the close proximity of unlearned parameters to the original model. To demonstrate their severity, we propose unlearning-specific membership inference and reconstruction attacks, showing that several state-of-the-art methods (e.g., NGP, SCRUB) remain vulnerable. To mitigate this leakage, we introduce WARP, a plug-and-play teleportation defense that leverages neural network symmetries to reduce forget-set gradient energy and increase parameter dispersion while preserving predictions. This reparameterization obfuscates the signal of forgotten data, making it harder for attackers to distinguish forgotten samples from non-members or recover them via reconstruction. Across six unlearning algorithms, our approach achieves consistent privacy gains, reducing adversarial advantage (AUC) by up to 64% in black-box and 92% in white-box settings, while maintaining accuracy on retained data. These results highlight teleportation as a general tool for reducing attack success in approximate unlearning.

📄 PDF Abstract BibTeX arXiv:2512.00272

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond

2025-02-07 · Chongyu Fan, Jinghan Jia, Yihua Zhang, Anil Ramakrishna 외

The LLM unlearning technique has recently been introduced to comply with data regulations and address the safety and ethical concerns of LLMs by removing the undesired data-model influence. However, state-of-the-art unle…

Neural Teleportation

2020-12-02 · Marco Armenta, Thierry Judge, Nathan Painchaud, Youssef Skandarani 외

In this paper, we explore a process called neural teleportation, a mathematical consequence of applying quiver representation theory to neural networks. Neural teleportation "teleports" a network to a new position in the…

Position

Do Unlearning Methods Remove Information from Language Model Weights?

2024-10-11 · Aghyad Deeb, Fabien Roger

Large Language Models' knowledge of how to perform cyber-security attacks, create bioweapons, and manipulate humans poses risks of misuse. Previous work has proposed methods to unlearn this knowledge. Historically, it ha…

Language ModelingLanguage Modelling

Machine Unlearning with Minimal Gradient Dependence for High Unlearning Ratios

2024-06-24 · Tao Huang, Ziyang Chen, Jiayang Meng, Qingyu Huang 외

In the context of machine unlearning, the primary challenge lies in effectively removing traces of private data from trained models while maintaining model performance and security against privacy attacks like membership…

Machine Unlearning

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs

2025-09-06 · Debdeep Sanyal, Manodeep Ray, Murari Mandal arxiv

The release of open-weight large language models (LLMs) creates a tension between advancing accessible research and preventing misuse, such as malicious fine-tuning to elicit harmful content. Current safety measures stru…