In Search of Grandmother Cells: Tracing Interpretable Neurons in Tabular Representations
Foundation models are powerful yet often opaque in their decision-making. A topic of continued interest in both neuroscience and artificial intelligence is whether some neurons behave like grandmother cells, i.e., neurons that are inherently interpretable because they exclusively respond to single concepts. In this work, we propose two information-theoretic measures that quantify the neuronal saliency and selectivity for single concepts. We apply these metrics to the representations of TabPFN, a tabular foundation model, and perform a simple search across neuron-concept pairs to find the most salient and selective pair. Our analysis provides the first evidence that some neurons in such models show moderate, statistically significant saliency and selectivity for high-level concepts. These findings suggest that interpretable neurons can emerge naturally and that they can, in some cases, be identified without resorting to more complex interpretability techniques.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
How do language models retrieve entity-specific facts from their parameters? We investigate this question by searching for sparse, entity-selective MLP neurons - which we call entity cells, by analogy to the "grandmother…
Prototype memory and attention mechanisms for few shot image generation
Recent discoveries indicate that the neural codes in the primary visual cortex (V1) of macaque monkeys are complex, diverse and sparse. This leads us to ponder the computational advantages and functional role of these “g…
Image GenerationOnline ClusteringClueReader: Heterogeneous Graph Attention Network for Multi-hop Machine Reading Comprehension
Multi-hop machine reading comprehension is a challenging task in natural language processing as it requires more reasoning ability across multiple documents. Spectral models based on graph convolutional networks have sho…
Graph AttentionMachine Reading ComprehensionReading ComprehensionLanguage Model Circuits Are Sparse in the Neuron Basis
The high-level concepts that a neural network uses to perform computation need not be aligned to individual neurons (Smolensky, 1986). Language model interpretability research has thus turned to techniques which decompos…
High-Precision Automated Reconstruction of Neurons with Flood-filling Networks
Reconstruction of neural circuits from volume electron microscopy data requires the tracing of complete cells including all their neurites. Automated approaches have been developed to perform the tracing, but without cos…
Electron Microscopy Image SegmentationImage ReconstructionVocal Bursts Intensity Prediction