paper-with-me

홈 › Papers

Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning

2026-05-08 · Chris Elliott, Daniel Murfet arxiv

These notes introduce the theory of susceptibilities as developed in [arXiv:2504.18274, arXiv:2601.12703] for interpreting neural networks. The susceptibility of an observable $φ$ to a data perturbation is defined as a derivative of a posterior expectation, which by the fluctuation--dissipation theorem equals a posterior covariance. Different choices of $φ$ yield different objects: per-sample losses give the influence matrix (the Bayesian influence function of [arXiv:2509.26544]), while component-localized observables give the structural susceptibility matrix that pairs model components with data patterns. The susceptibility matrix is (up to a factor of $nβ$) the Jacobian of the map from data distributions to structural coordinates; its pseudo-inverse provides a linearized solution to the patterning problem of [arXiv:2601.13548]: finding data perturbations that produce a desired structural change. We motivate the theory from its statistical-mechanical foundations, then give a detailed exposition of susceptibilities, their empirical estimators, and their connection to the geometry of the loss landscape.

📄 PDF Abstract BibTeX arXiv:2605.07980

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Patterning: The Dual of Interpretability

2026-01-20 · George Wang, Daniel Murfet arxiv

Mechanistic interpretability aims to understand how neural networks generalize beyond their training data by reverse-engineering their internal structures. We introduce patterning as the dual problem: given a desired for…

Structural Inference: Interpreting Small Language Models with Susceptibilities

2025-04-25 · Garrett Baker, George Wang, Jesse Hoogland, Daniel Murfet

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward Gi…

Linear Response Estimators for Singular Statistical Models

2026-05-08 · Chris Elliott, Daniel Murfet arxiv

We define susceptibilities as a measure of the response of an observable quantity of a parameterized statistical model to a perturbation of the data for a general class of observables. We define estimators for these susc…

Interpreting Reinforcement Learning Agents with Susceptibilities

2026-05-08 · Chris Elliott, Einar Urdshals, David Quarel, Daniel Murfet arxiv

Susceptibilities are a technique for neural network interpretability that studies the response of posterior expectation values of observables to perturbations of the loss. We generalize this construction to the setting o…

Reinforcement Learning

Towards Spectroscopy: Susceptibility Clusters in Language Models

2026-01-19 · Andrew Gordon, Garrett Baker, George Wang, William Snell 외 arxiv

Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distribution by upweighting a token $y$ in cont…