paper-with-me

Papers

Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution

2026-05-14 · Saisab Sadhu, Pratinav Seth, Vinay Kumar Sankarapu arxiv

Standard unlearning evaluations measure behavioral suppression in full precision, immediately after training, despite every deployed language model being quantized first. Recent work has shown that 4-bit post-training quantization can reverse machine unlearning; we show this is not a tuning artefact but a systematic dual failure: gradient-based methods that achieve meaningful forgetting lose it under compression, while methods that survive quantization barely change the model. Both failures trace to the same root cause: across all baselines, per-parameter updates lie 47-828x below the NF4 quantization bin width; updates diffused across billions of parameters cannot clear quantization bin boundaries, a consequence we formalize as a sparsity-permanence tradeoff. We present MANSU (Mechanistic-Aligned Null-Space Unlearning), which resolves both modes by combining causal circuit attribution to isolate the minimal forget-set subgraph, circuit-restricted null-space projection with a diagonal-Fisher retain bound, and a per-parameter magnitude floor guaranteeing quantization survival by construction. We additionally introduce Circuit Attribution Divergence (CAD), a mechanistic verification metric distinguishing structural erasure from behavioral suppression, a distinction existing metrics cannot make. Across multiple model families and hazard benchmarks, MANSU is the first method to jointly satisfy all four properties with margin on each (meaningful forgetting, retain preservation, non-positive PTQ gap, and structural erasure), while gradient-based baselines recover up to +0.05 accuracy under compression.

📄 PDF Abstract BibTeX arXiv:2605.15138

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

2026-09-09 · Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou arxiv

Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers a practical alternat…

QUAIL: Quantization Aware Unlearning for Mitigating Misinformation in LLMs

2026-01-21 · Himanshu Mishra, Kanwal Mehreen arxiv

Machine unlearning aims to remove specific knowledge (e.g., copyrighted or private data) from a trained model without full retraining. In practice, models are often quantized (e.g., 4-bit) for deployment, but we find tha…

Quantization-Robust LLM Unlearning via Low-Rank Adaptation

2026-02-13 · João Vitor Boer Abitante, Joana Meneguzzo Pasquali, Luan Fonseca Garcia, Ewerton de Oliveira 외 arxiv

Large Language Model (LLM) unlearning aims to remove targeted knowledge from a trained model, but practical deployments often require post-training quantization (PTQ) for efficient inference. However, aggressive low-bit …

Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting

2025-11-13 · James Jin Kang, Dang Bui, Thanh Pham, Huo-Chong Ling arxiv

The growing use of large language models in sensitive domains has exposed a critical weakness: the inability to ensure that private information can be permanently forgotten. Yet these systems still lack reliable mechanis…

Federated Learning

DurableUn: Quantization-Induced Recovery Attacks in Machine Unlearning

2026-05-04 · Abdullah Ahmad Khan, Ferdous Sohel arxiv

Machine unlearning aims to remove specified training data to satisfy privacy regulations such as GDPR. However, existing evaluations assume identical precision at unlearning and deployment, overlooking that production LL…