paper-with-me

홈 › Papers

Neuron Empirical Gradient: Connecting Neurons' Linear Controllability and Representational Capacity

2024-12-24 · Xin Zhao, Zehui Jiang, Naoki Yoshinaga

Although neurons in the feed-forward layers of pre-trained language models (PLMs) can store factual knowledge, most prior analyses remain qualitative, leaving the quantitative relationship among knowledge representation, neuron activations, and model output poorly understood. In this study, by performing neuron-wise interventions using factual probing datasets, we first reveal the linear relationship between neuron activations and output token probabilities. We refer to the gradient of this linear relationship as ``neuron empirical gradients.'' and propose NeurGrad, an efficient method for their calculation to facilitate quantitative neuron analysis. We next investigate whether neuron empirical gradients in PLMs encode general task knowledge by probing skill neurons. To this end, we introduce MCEval8k, a multi-choice knowledge evaluation benchmark spanning six genres and 22 tasks. Our experiments confirm that neuron empirical gradients effectively capture knowledge, while skill neurons exhibit efficiency, generality, inclusivity, and interdependency. These findings link knowledge to PLM outputs via neuron empirical gradients, shedding light on how PLMs store knowledge. The code and dataset are released.

📄 PDF Abstract BibTeX arXiv:2412.18053

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Connecting Spiking Neurons to a Spiking Memristor Network Changes the Memristor Dynamics

2014-02-17 · Deborah Gater, Attya Iqbal, Jeffrey Davey, Ella Gale

Memristors have been suggested as neuromorphic computing elements. Spike-time dependent plasticity and the Hodgkin-Huxley model of the neuron have both been modelled effectively by memristor theory. The d.c. response of …

Translation

MgNO: Efficient Parameterization of Linear Operators via Multigrid

2023-10-16 · Juncai He, Xinliang Liu, Jinchao Xu

In this work, we propose a concise neural operator architecture for operator learning. Drawing an analogy with a conventional fully connected neural network, we define the neural operator as follows: the output of the $i…

Operator learning

SpaRCe: Improved Learning of Reservoir Computing Systems through Sparse Representations

2019-12-04 · Luca Manneschi, Andrew C. Lin, Eleni Vasilaki

"Sparse" neural networks, in which relatively few neurons or connections are active, are common in both machine learning and neuroscience. Whereas in machine learning, "sparsity" is related to a penalty term that leads t…

BIG-bench Machine LearningDecision Making

Training capsules as a routing-weighted product of expert neurons

2019-07-26 · Michael Hauser

Capsules are the multidimensional analogue to scalar neurons in neural networks, and because they are multidimensional, much more complex routing schemes can be used to pass information forward through the network than w…

The Mechanical Neural Network(MNN) -- A physical implementation of a multilayer perceptron for education and hands-on experimentation

2022-07-15 · Axel Schaffland

In this paper the Mechanical Neural Network(MNN) is introduced, a physical implementation of a multilayer perceptron(MLP) with ReLU activation functions, two input neurons, four hidden neurons and two output neurons. Thi…