paper-with-me

홈 › Papers

From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability

2025-11-17 · Jihoon Moon arxiv

Deep neural networks achieve state of the art performance but remain difficult to interpret mechanistically. In this work, we propose a control theoretic framework that treats a trained neural network as a nonlinear state space system and uses local linearization, controllability and observability Gramians, and Hankel singular values to analyze its internal computation. For a given input, we linearize the network around the corresponding hidden activation pattern and construct a state space model whose state consists of hidden neuron activations. The input state and state output Jacobians define local controllability and observability Gramians, from which we compute Hankel singular values and associated modes. These quantities provide a principled notion of neuron and pathway importance: controllability measures how easily each neuron can be excited by input perturbations, observability measures how strongly each neuron influences the output, and Hankel singular values rank internal modes that carry input output energy. We illustrate the framework on simple feedforward networks, including a 1 2 2 1 SwiGLU network and a 2 3 3 2 GELU network. By comparing different operating points, we show how activation saturation reduces controllability, shrinks the dominant Hankel singular value, and shifts the dominant internal mode to a different subset of neurons. The proposed method turns a neural network into a collection of local white box dynamical models and suggests which internal directions are natural candidates for pruning or constraints to improve interpretability.

📄 PDF Abstract BibTeX arXiv:2511.12852

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unlocking Interpretability for RF Sensing: A Complex-Valued White-Box Transformer

2025-07-29 · Xie Zhang, Yina Wang, Chenshu Wu arxiv

The empirical success of deep learning has spurred its application to the radio-frequency (RF) domain, leading to significant advances in Deep Wireless Sensing (DWS). However, most existing DWS models function as black b…

Steered LLM Activations are Non-Surjective

2026-04-10 · Aayush Mishra, Daniel Khashabi, Anqi Liu arxiv

Activation steering is a popular white-box control technique that modifies model activations to elicit an abstract change in its behavior. It has also become a standard tool in interpretability (e.g., probing truthfulnes…

Discovering Continuous-Time Memory-Based Symbolic Policies using Genetic Programming

2024-06-04 · Sigur de Vries, Sander Keemink, Marcel van Gerven

Artificial intelligence techniques are increasingly being applied to solve control problems, but often rely on black-box methods without transparent output generation. To improve the interpretability and transparency in …

Evaluating Attribution Methods using White-Box LSTMs

2020-10-16 · EMNLP (BlackboxNLP) 2020 11 · Yiding Hao

Interpretability methods for neural networks are difficult to evaluate because we do not understand the black-box models typically used to test them. This paper proposes a framework in which interpretability methods are …

Integrating White and Black Box Techniques for Interpretable Machine Learning

2024-07-12 · Eric M. Vernon, Naoki Masuyama, Yusuke Nojima

In machine learning algorithm design, there exists a trade-off between the interpretability and performance of the algorithm. In general, algorithms which are simpler and easier for humans to comprehend tend to show wors…

Interpretable Machine Learning