paper-with-me

홈 › Papers

On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning

2026-05-26 · Xiaotian Ye, Xiaohan Wang, Mengqi Zhang, Shu Wu arxiv

Counterfactual tuning (CFT) has emerged as a promising paradigm for Large Language Model (LLM) unlearning by training models to generate alternative fictitious knowledge in place of undesired content. However, in this work, we find that this paradigm still underperforms other paradigms in some aspects, and identify two previously overlooked pitfalls underlying this gap: (1) knowledge conflict, where mutual inconsistencies within counterfactual corpora induce conflicting gradients that disrupt parameter optimization, and (2) hallucination spillover, where fitting false targets instills a persistent fabrication bias, inflating hallucination rates on unrelated domains. To systematically diagnose these issues, we introduce RWKU+, an extended benchmark equipped with novel trade-off metrics and gradient-level diagnostic tools. Our work further discusses the limitations and overhead of the paradigm, aiming to provide insights and actionable guidance for more rigorous LLM unlearning research.

📄 PDF Abstract BibTeX arXiv:2605.27083

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Probing Knowledge Holes in Unlearned LLMs

2025-10-27 · Myeongseob Ko, Hoang Anh Just, Charles Fleming, Ming Jin 외 arxiv

Machine unlearning has emerged as a prevalent technical solution for selectively removing unwanted knowledge absorbed during pre-training, without requiring full retraining. While recent unlearning techniques can effecti…

Unlearning through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse Adapter

2025-01-01 · CVPR 2025 1 · Zhengyi Zhong, Weidong Bao, Ji Wang, Shuai Zhang 외

Federated Learning is a promising paradigm for privacy-preserving collaborative model training. In practice, it is essential not only to continuously train the model to acquire new knowledge but also to guarantee old…

Federated LearningPrivacy Preserving

CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization

2026-04-17 · Junyi Li, Yongqiang Chen, Ningning Ding arxiv

Machine unlearning has gained increasing attention in recent years, as a promising technique to selectively remove unwanted privacy or copyrighted information from Large Language Models that are trained on a massive scal…

Debiasing Machine Unlearning with Counterfactual Examples

2024-04-24 · Ziheng Chen, Jia Wang, Jun Zhuang, Abbavaram Gowtham Reddy 외

The right to be forgotten (RTBF) seeks to safeguard individuals from the enduring effects of their historical actions by implementing machine-learning techniques. These techniques facilitate the deletion of previously ac…

counterfactualMachine Unlearning

GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning

2026-01-30 · Naoki Murata, Yuhta Takida, Chieh-Hsin Lai, Toshimitsu Uesaka 외 arxiv

Training-data attribution for vision generative models aims to identify which training data influenced a given output. While most methods score individual examples, practitioners often need group-level answers (e.g., art…

Semantic Similarity