paper-with-me

홈 › Papers

The Limits of Obliviate: Evaluating Unlearning in LLMs via Stimulus-Knowledge Entanglement-Behavior Framework

2025-10-29 · Aakriti Shah, Thai Le arxiv

Unlearning in large language models (LLMs) is crucial for managing sensitive data and correcting misinformation, yet evaluating its effectiveness remains an open problem. We investigate whether persuasive prompting can recall factual knowledge from deliberately unlearned LLMs across models ranging from 2.7B to 13B parameters (OPT-2.7B, LLaMA-2-7B, LLaMA-3.1-8B, LLaMA-2-13B). Drawing from ACT-R and Hebbian theory (spreading activation theories), as well as communication principles, we introduce Stimulus-Knowledge Entanglement-Behavior Framework (SKeB), which models information entanglement via domain graphs and tests whether factual recall in unlearned models is correlated with persuasive framing. We develop entanglement metrics to quantify knowledge activation patterns and evaluate factuality, non-factuality, and hallucination in outputs. Our results show persuasive prompts substantially enhance factual knowledge recall (14.8% baseline vs. 24.5% with authority framing), with effectiveness inversely correlated to model size (128% recovery in 2.7B vs. 15% in 13B). SKeB provides a foundation for assessing unlearning completeness, robustness, and overall behavior in LLMs.

📄 PDF Abstract BibTeX arXiv:2510.25732

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models

2025-02-20 · Mark Russinovich, Ahmed Salem

Recent copyright agreements between AI companies and content creators underscore the need for fine-grained control over language models' ability to reproduce copyrighted text. Existing defenses-ranging from aggressive un…

HellaSwagMemorizationMMLUTruthfulQA+1

OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models

2025-05-07 · Xiaoyu Xu, Minxin Du, Qingqing Ye, Haibo Hu

Large language models (LLMs) trained over extensive corpora risk memorizing sensitive, copyrighted, or toxic content. To address this, we propose OBLIVIATE, a robust unlearning framework that removes targeted data while …

Machine UnlearningMemorization

DeepObliviate: A Powerful Charm for Erasing Data Residual Memory in Deep Neural Networks

2021-05-13 · Yingzhe He, Guozhu Meng, Kai Chen, Jinwen He 외

Machine unlearning has great significance in guaranteeing model security and protecting user privacy. Additionally, many legal provisions clearly stipulate that users have the right to demand model providers to delete th…

Machine Unlearning

Analyzing and Mitigating Object Hallucination: A Training Bias Perspective

2025-08-06 · Yifan Li, Kun Zhou, Wayne Xin Zhao, Lei Fang 외 arxiv

As scaling up training data has significantly improved the general multimodal capabilities of Large Vision-Language Models (LVLMs), they still suffer from the hallucination issue, generating text that is inconsistent wit…

Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting

2025-11-13 · James Jin Kang, Dang Bui, Thanh Pham, Huo-Chong Ling arxiv

The growing use of large language models in sensitive domains has exposed a critical weakness: the inability to ensure that private information can be permanently forgotten. Yet these systems still lack reliable mechanis…

Federated Learning