paper-with-me

홈 › Papers

Interpreting Reinforcement Learning Agents with Susceptibilities

2026-05-08 · Chris Elliott, Einar Urdshals, David Quarel, Daniel Murfet arxiv

Susceptibilities are a technique for neural network interpretability that studies the response of posterior expectation values of observables to perturbations of the loss. We generalize this construction to the setting of the regret in deep reinforcement learning and investigate the utility of susceptibilities in a simple gridworld model that nevertheless exhibits non-trivial stagewise development. We argue that susceptibilities reveal internal features of the development of the model in parameter space that one cannot detect purely by studying the development of the learned policy. We validate these results with activation-steering, and discuss the framework's extension to RLHF post-training.

📄 PDF Abstract BibTeX arXiv:2605.08007

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Structural Inference: Interpreting Small Language Models with Susceptibilities

2025-04-25 · Garrett Baker, George Wang, Jesse Hoogland, Daniel Murfet

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward Gi…

Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning

2026-05-08 · Chris Elliott, Daniel Murfet arxiv

These notes introduce the theory of susceptibilities as developed in [arXiv:2504.18274, arXiv:2601.12703] for interpreting neural networks. The susceptibility of an observable $φ$ to a data perturbation is defined as a d…

Linear Response Estimators for Singular Statistical Models

2026-05-08 · Chris Elliott, Daniel Murfet arxiv

We define susceptibilities as a measure of the response of an observable quantity of a parameterized statistical model to a perturbation of the data for a general class of observables. We define estimators for these susc…

PARL: Prompt-based Agents for Reinforcement Learning

2025-10-24 · Yarik Menchaca Resendiz, Roman Klinger arxiv

Large language models (LLMs) have demonstrated high performance on tasks expressed in natural language, particularly in zero- or few-shot settings. These are typically framed as supervised (e.g., classification) or unsup…

Reinforcement Learning

Generating Black-Box Adversarial Examples for Text Classifiers Using a Deep Reinforced Model

2019-09-17 · Prashanth Vijayaraghavan, Deb Roy

Recently, generating adversarial examples has become an important means of measuring robustness of a deep learning model. Adversarial examples help us identify the susceptibilities of the model and further counter those …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Sentiment Analysis+1