paper-with-me

홈 › Papers

Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial Editing

2025-05-19 · Jiakuan Xie, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao

Knowledge editing, which aims to update the knowledge encoded in language models, can be deceptive. Despite the fact that many existing knowledge editing algorithms achieve near-perfect performance on conventional metrics, the models edited by them are still prone to generating original knowledge. This paper introduces the concept of "superficial editing" to describe this phenomenon. Our comprehensive evaluation reveals that this issue presents a significant challenge to existing algorithms. Through systematic investigation, we identify and validate two key factors contributing to this issue: (1) the residual stream at the last subject position in earlier layers and (2) specific attention modules in later layers. Notably, certain attention heads in later layers, along with specific left singular vectors in their output matrices, encapsulate the original knowledge and exhibit a causal relationship with superficial editing. Furthermore, we extend our analysis to the task of superficial unlearning, where we observe consistent patterns in the behavior of specific attention heads and their corresponding left singular vectors, thereby demonstrating the robustness and broader applicability of our methodology and conclusions. Our code is available here.

📄 PDF Abstract BibTeX arXiv:2505.12636

Code (1)

jiakuan929/superficial-editing 공식 구현 pytorch

Tasks

knowledge editing

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Exposing the Illusion of Erasure in Knowledge Editing for LLMs

2026-06-22 · Advik Raj Basani, Anshuman Chhabra arxiv

Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an …

knowledge editing

On Mechanistic Knowledge Localization in Text-to-Image Generative Models

2024-05-02 · Samyadeep Basu, Keivan Rezaei, Priyatham Kattakinda, Ryan Rossi 외

Identifying layers within text-to-image models which control visual attributes can facilitate efficient model editing through closed-form updates. Recent work, leveraging causal tracing show that early Stable-Diffusion v…

Model Editing

Addressing the Reasoning Gap: Mechanistic Circuit-Based Knowledge Editing in Large Language Models

2026-04-07 · Tianyi Zhao, Yinhan He, Wendy Zheng, Chen Chen arxiv

Deploying Large Language Models (LLMs) in real-world dynamic environments raises the challenge of updating their pre-trained knowledge. While existing knowledge editing methods can reliably patch isolated facts, they fre…

knowledge editing

Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization

2024-10-16 · Phillip Guo, Aaquib Syed, Abhay Sheshadri, Aidan Ewart 외

Methods for knowledge editing and unlearning in large language models seek to edit or remove undesirable knowledge or capabilities without compromising general language modeling performance. This work investigates how me…

knowledge editingLanguage ModelingLanguage Modelling

LMs as Task-Specific Knowledge Bases: An Interpretability Analysis

2026-06-25 · Amit Elhelo, Amir Globerson, Mor Geva arxiv

Language models (LMs) capture large amounts of factual knowledge applicable to a wide range of tasks, motivating the view of their parameters as a knowledge base. An important property of knowledge bases is that differen…