paper-with-me

Papers

IG-Lens: Exact Additive Probability Attribution Across Transformer Layers via Telescoping Integrated Gradients

2026-06-29 · Duc Anh Nguyen arxiv

We ask a simple question about decoder-only transformers: between which two layers is the probability of a predicted token actually produced? Existing layer-wise readout tools answer only approximately. The logit lens and its trained variant report a per-layer level of probability but give no additive decomposition; their estimates are biased and non-monotone across depth. Direct Logit Attribution and related residual-stream methods are additive, but only in logit space, the softmax nonlinearity breaks additivity in probability space, precisely the quantity one usually cares about. Layer Conductance integrates gradients per layer, but attributes each to its own baseline and so does not sum to the total change in prediction. We introduce IG-Lens, a telescoping application of Integrated Gradients along a single path through the hidden states from a baseline to the final layer. Crediting each segment to the layer it terminates at yields a layer-wise attribution whose sum is exactly the change in target probability, with the softmax inside the integration path rather than linearized away. Our default estimator credits each integration step its observed change in target probability (a prediction-aware reweighting in the spirit of IDGI) rather than its raw gradient. Because the readout is a one-dimensional probability, this collapses each segment to a telescoping sum of endpoint values, so completeness holds exactly (to floating point) at any step count, removing Riemann discretization error while suppressing steps that show gradient sensitivity without a change in output. We give the telescoping identity and its proof, verify completeness to floating point, and describe a single-pass batched implementation computing the full token-by-layer map without any backward call. Code: https://github.com/anhnda/IGLens.

📄 PDF Abstract BibTeX arXiv:2606.29693

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MolLedger: An Additive Graph Neural Network with Chemically Grounded ADME Attributions

2026-08-31 · Christina X. Ji arxiv

Optimizing absorption, distribution, metabolism, and excretion (ADME) is an important part of small molecule drug discovery. Many machine learning models have been built to predict ADME properties to facilitate this opti…

Graph Neural NetworkDrug Discovery

A Polynomial Architecture-Attribution Co-Design Framework for Exact Aumann-Shapley Attribution in GNNs

2026-07-23 · Bizu Feng, Zhimu Yang, Shuming Wang, Shaode Yu 외 arxiv

We study feature-level and node-level explanations for graph neural networks (GNNs) through the lens of Aumann-Shapley attribution. Path-integral methods such as Integrated Gradients provide an axiomatic formulation of a…

Rethinking Log Odds: Linear Probability Modelling and Expert Advice in Interpretable Machine Learning

2022-11-11 · Danial Dervovic, Nicolas Marchesotti, Freddy Lecue, Daniele Magazzeni

We introduce a family of interpretable machine learning models, with two broad additions: Linearised Additive Models (LAMs) which replace the ubiquitous logistic link function in General Additive Models (GAMs); and Subsc…

Additive modelsBinary ClassificationInterpretable Machine Learning

Consistent feature attribution for tree ensembles

2017-06-19 · Scott M. Lundberg, Su-In Lee

Note that a newer expanded version of this paper is now available at: arXiv:1802.03888 It is critical in many applications to understand what features are important for a model, and why individual predictions were made…

ClusteringFeature Importance

How Faithful Is Attribution for Sales Forecasting? A Counterfactual Study

2026-09-04 · Glib Kechyn arxiv

Deep models for sales forecasting, such as WaveNet-style dilated convolutional networks, are accurate but opaque: when a single model predicts sales for one of many series, it offers no account of why. We add a post-hoc,…