paper-with-me

홈 › Papers

Targeted Neuron Modulation via Contrastive Pair Search

2026-05-12 · Sam Herring, Jake Naviasky, Karan Malhotra arxiv

Language models are instruction-tuned to refuse harmful requests, but the mechanisms underlying this behavior remain poorly understood. Popular steering methods operate on the residual stream and degrade output coherence at high intervention strengths, limiting their practical use. We introduce contrastive neuron attribution (CNA), which identifies the 0.1% of MLP neurons whose activations most distinguish harmful from benign prompts, requiring only forward passes with no gradients or auxiliary training. In instruct models, ablating the discovered circuit reduces refusal rates by over 50% on a standard jailbreak benchmark while preserving fluency and non-degeneracy across all steering strengths. Applying CNA to matched base and instruct models across Llama and Qwen architectures (from 1B to 72B parameters), we find that base models contain similar late-layer discrimination structures but steering these neurons produces only content shifts, not behavioral change. These results demonstrate that neuron-level intervention enables reliable behavioral steering without the quality tradeoffs of residual-stream methods. More broadly, our findings suggest that alignment fine-tuning transforms pre-existing discrimination structure into a sparse, targetable refusal gate.

📄 PDF Abstract BibTeX arXiv:2605.12290

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Flexible information routing in neural populations through stochastic comodulation

2019-12-01 · NeurIPS 2019 12 · Caroline Haimerl, Cristina Savin, Eero Simoncelli

Humans and animals are capable of flexibly switching between a multitude of tasks, each requiring rapid, sensory-informed decision making. Incoming stimuli are processed by a hierarchy of neural circuits consisting of mi…

Decision MakingDecoder

Neuromodulation and homeostasis: complementary mechanisms for robust neural function

2024-12-05 · Arthur Fyon, Guillaume Drion

Neurons depend on two interdependent mechanisms-homeostasis and neuromodulation-to maintain robust and adaptable functionality. Homeostasis stabilizes neuronal activity by adjusting ionic conductances, whereas neuromodul…

Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents

2024-09-25 · Emanuela Boros, Maud Ehrmann

This paper investigates the presence of OCR-sensitive neurons within the Transformer architecture and their influence on named entity recognition (NER) performance on historical documents. By analysing neuron activation …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Noninvasive precision modulation of high-level neural population activity via natural vision perturbations

2025-06-05 · Guy Gaziv, Sarah Goulding, Ani Ayvazian-Hancock, Yoon Bai 외

Precise control of neural activity -- modulating target neurons deep in the brain while leaving nearby neurons unaffected -- is an outstanding challenge in neuroscience, generally approached using invasive techniques. Th…

Neuron Platonic Intrinsic Representation From Dynamics Using Contrastive Learning

2025-02-06 · Wei Wu, Can Liao, Zizhen Deng, Zhengrui Guo 외

The Platonic Representation Hypothesis suggests a universal, modality-independent reality representation behind different data modalities. Inspired by this, we view each neuron as a system and detect its multi-segment ac…

Contrastive Learning