From Neurons to Neutrons: A Case Study in Interpretability
Mechanistic Interpretability (MI) promises a path toward fully understanding how neural networks make their predictions. Prior work demonstrates that even when trained to perform simple arithmetic, models can implement a variety of algorithms (sometimes concurrently) depending on initialization and hyperparameters. Does this mean neuron-level interpretability techniques have limited applicability? We argue that high-dimensional neural networks can learn low-dimensional representations of their training data that are useful beyond simply making good predictions. Such representations can be understood through the mechanistic interpretability lens and provide insights that are surprisingly faithful to human-derived domain knowledge. This indicates that such approaches to interpretability can be useful for deriving a new understanding of a problem from models trained to solve it. As a case study, we extract nuclear physics concepts by studying models trained to reproduce nuclear data.
Code (1)
Similar Papers 제목 키워드 기반
Irradiation Tests for Commercial Off-the Shelf Components with Atmospheric-like Neutrons and Heavy-Ions
This paper presents the results of the irradiation, performed with atmospheric-like neutrons and heavy-ions, of Commercial Off-the Shelf Components (COTS), which can be used in space missions. In such cases, it is crucia…
NeutronStream: A Dynamic GNN Training Framework with Sliding Window for Graph Streams
Existing Graph Neural Network (GNN) training frameworks have been designed to help developers easily create performant GNN implementations. However, most existing GNN frameworks assume that the input graphs are static, b…
Graph Neural NetworkA Case Study on Concept Induction for Neuron-Level Interpretability in CNN
Deep Neural Networks (DNNs) have advanced applications in domains such as healthcare, autonomous systems, and scene understanding, yet the internal semantics of their hidden neurons remain poorly understood. Prior work i…
Scene UnderstandingScene RecognitionAdjusting the nuclear reactor's neutron transport and diffusion theory for an alternative description and modelling of postage or supplies delivery processes
There seems to exist significant similarities between a reactor system and a supply chain from collection to delivery. In the reactor case, neutrons are continuously produced and absorbed in nuclear fuel. In a supply sys…
UnityImprovement studies on neutron-gamma separation in HPGe detectors by using neural networks
The neutrons emitted in heavy-ion fusion-evaporation (HIFE) reactions together with the gamma-rays cause unwanted backgrounds in gamma-ray spectra. Especially in the nuclear reactions, where relativistic ion beams (RIBs)…