Leverage Is Not Reach: A Control-Window Law for Single-Neuron Steering in Language Models
Aligned language models gate behaviors such as refusal and language routing through sparse feed forward neurons, yet no theory predicts when a single neuron intervention controls a behavior coherently rather than collapsing the output. We develop a budget normalized control window framework for single neuron steering. A dose along one write direction reduces to one control coordinate: the alignment between the residual stream and the write, driven along a universal saturation curve in units of a coherence budget set by the residual norm divided by the write norm. Coherent control exists when a behavior trigger lies below the collapse ceiling. The same coordinate governs benign mode switches and refusal; the ceiling follows from weights and one generic forward pass, while triggers are measured at rollout. On fifteen held out neurons, the predicted ceiling has mean absolute error 0.14, about 0.07 in bulk layers, and the committed open or closed verdict holds on eleven against a ten of fifteen majority baseline. Closed cases expose three failure modes rather than violations: collapse before trigger, too little depth to propagate, or a normalization that caps how far one neuron can push. The law explains why local gradient attribution anti predicts control: true controllers write off the readout axis and carry a near zero first order gradient. A forward only contrastive screen made precise by the window recovers controllers that attribution misses. On refusal, the hardest case, intervention success is typed, not scalar: coherent bypass and strict actionable reach separate, so a neuron can flip refusal in fluent, on task text with no actionable content, and genuine actionable reach appears only for three of six audited Llama pivots and only at later rollout horizons. Single neuron steering is therefore a budgeted, typed audit of controllability rather than a fixed dose anecdote.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Analysis of a Simple Neuromorphic Controller for Linear Systems: A Hybrid Systems Perspective
In this paper we analyze a neuromorphic controller, inspired by the leaky integrate-and-fire neuronal model, in closed-loop with a single-input single-output linear time-invariant system. The controller consists of two n…
Functional Cliques in Developmentally Correlated Neural Networks
We consider a sparse random network of excitatory leaky integrate-and-fire neurons with short-term synaptic depression. Furthermore to mimic the dynamics of a brain circuit in its first stages of development we introduce…
Reaching Optimized Parameter Set, Protein Secondary Structure Prediction Using Neural Network
We propose an optimized parameter set for protein secondary structure prediction using three layer feed forward back propagation neural network. The methodology uses four parameters viz. encoding scheme, window size, num…
Protein Secondary Structure PredictionSpecificitySSI-GAN: Semi-Supervised Swin-Inspired Generative Adversarial Networks for Neuronal Spike Classification
Mosquitos are the main transmissive agents of arboviral diseases. Manual classification of their neuronal spike patterns is very labor-intensive and expensive. Most available deep learning solutions require fully labeled…
Efficient Architecture Search for Continual Learning
Continual learning with neural networks is an important learning framework in AI that aims to learn a sequence of tasks well. However, it is often confronted with three challenges: (1) overcome the catastrophic forgettin…
Continual LearningNeural Architecture SearchTransfer Learning