paper-with-me

홈 › Papers

Journey to the Center of the Knowledge Neurons: Discoveries of Language-Independent Knowledge Neurons and Degenerate Knowledge Neurons

2023-08-25 · YuHeng Chen, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao

Pre-trained language models (PLMs) contain vast amounts of factual knowledge, but how the knowledge is stored in the parameters remains unclear. This paper delves into the complex task of understanding how factual knowledge is stored in multilingual PLMs, and introduces the Architecture-adapted Multilingual Integrated Gradients method, which successfully localizes knowledge neurons more precisely compared to current methods, and is more universal across various architectures and languages. Moreover, we conduct an in-depth exploration of knowledge neurons, leading to the following two important discoveries: (1) The discovery of Language-Independent Knowledge Neurons, which store factual knowledge in a form that transcends language. We design cross-lingual knowledge editing experiments, demonstrating that the PLMs can accomplish this task based on language-independent neurons; (2) The discovery of Degenerate Knowledge Neurons, a novel type of neuron showing that different knowledge neurons can store the same fact. Its property of functional overlap endows the PLMs with a robust mastery of factual knowledge. We design fact-checking experiments, proving that the degenerate knowledge neurons can help the PLMs to detect wrong facts. Experiments corroborate these findings, shedding light on the mechanisms of factual knowledge storage in multilingual PLMs, and contribute valuable insights to the field. The code is available at https://github.com/heng840/AMIG.

📄 PDF Abstract BibTeX arXiv:2308.13198

Code (1)

heng840/amig 공식 구현 pytorch

Tasks

Fact Checkingknowledge editing

Similar Papers 제목 키워드 기반

One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging

2026-04-03 · Baban Gain, Asif Ekbal, Trilok Nath Singh arxiv

Weight-space model merging combines independently fine-tuned models without accessing original training data, offering a practical alternative to joint training. While merging succeeds in multitask settings, its behavior…

Machine Translation

ChatGPT Alternative Solutions: Large Language Models Survey

2024-03-21 · Hanieh Alipour, Nick Pendar, Kohinoor Roy

In recent times, the grandeur of Large Language Models (LLMs) has not only shone in the realm of natural language processing but has also cast its brilliance across a vast array of applications. This remarkable display o…

BenchmarkingChatbotSurvey

Center-Embedding and Constituency in the Brain and a New Characterization of Context-Free Languages

2022-06-27 · Daniel Mitropolsky, Adiba Ejaz, Mirah Shi, Mihalis Yannakakis 외

A computational system implemented exclusively through the spiking of neurons was recently shown capable of syntax, that is, of carrying out the dependency parsing of simple English sentences. We address two of the most …

Dependency ParsingSentence

Exposing the Functionalities of Neurons for Gated Recurrent Unit Based Sequence-to-Sequence Model

2023-03-27 · Yi-Ting Lee, Da-Yi Wu, Chih-Chun Yang, Shou-De Lin

The goal of this paper is to report certain scientific discoveries about a Seq2Seq model. It is known that analyzing the behavior of RNN-based models at the neuron level is considered a more challenging task than analyzi…

Position

Knowledge Graph and Hypergraph Transformers with Repository-Attention and Journey-Based Role Transport

2026-02-07 · Mahesh Godavarti arxiv

We present a concise architecture for joint training on sentences and structured data while keeping knowledge and language representations separable. The model treats knowledge graphs and hypergraphs as structured instan…

Knowledge GraphsLink Prediction