paper-with-me

홈 › Papers

Understanding Language Model Circuits through Knowledge Editing

2024-06-25 · Huaizhi Ge, Frank Rudzicz, Zining Zhu

Recent advances in language model interpretability have identified circuits, critical subnetworks that replicate model behaviors, yet how knowledge is structured within these crucial subnetworks remains opaque. To gain an understanding toward the knowledge in the circuits, we conduct systematic knowledge editing experiments on the circuits of the GPT-2 language model. Our analysis reveals intriguing patterns in how circuits respond to editing attempts, the extent of knowledge distribution across network components, and the architectural composition of knowledge-bearing circuits. These findings offer insights into the complex relationship between model circuits and knowledge representation, deepening the understanding of how information is organized within language models. Our findings offer novel insights into the ``meanings'' of the circuits, and introduce directions for further interpretability and safety research of language models.

📄 PDF Abstract BibTeX arXiv:2406.17241

Code (0)

등록된 구현이 없습니다.

Tasks

knowledge editingLanguage ModelingLanguage Modellingmodeltext-classificationText Classification

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Multi-Head Attention 설명 없음
Weight Decay 설명 없음
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…

Similar Papers 제목 키워드 기반

Knowledge Circuits in Pretrained Transformers

2024-05-28 · Yunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang 외

The remarkable capabilities of modern large language models are rooted in their vast repositories of knowledge encoded within their parameters, enabling them to perceive the world and engage in reasoning. The inner worki…

In-Context Learningknowledge editingLanguage ModelingLanguage Modelling

CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners

2025-03-20 · Yunzhi Yao, Jizhan Fang, Jia-Chen Gu, Ningyu Zhang 외

Knowledge Editing (KE) enables the modification of outdated or incorrect information in large language models (LLMs). While existing KE methods can update isolated facts, they struggle to generalize these updates to mult…

knowledge editing

How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training

2025-02-16 · Yixin Ou, Yunzhi Yao, Ningyu Zhang, Hui Jin 외

Despite exceptional capabilities in knowledge-intensive tasks, Large Language Models (LLMs) face a critical gap in understanding how they internalize new knowledge, particularly how to structurally embed acquired knowled…

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

2024-03-28 · Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov 외

We introduce methods for discovering and applying sparse feature circuits. These are causally implicated subnetworks of human-interpretable features for explaining language model behaviors. Circuits identified in prior w…

Language ModelingLanguage Modelling

Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs

2026-04-20 · Tingchao Fu, Wenkai Wang, Fanxiao Li, Huadong Zhang 외 arxiv

Although Knowledge Editing provides an efficient mechanism for updating the knowledge of Multimodal Large Language Models (MLLMs), we find that current paradigms still suffer from an important yet remain underexplored is…

knowledge editing