paper-with-me

홈 › Papers

nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers

2025-11-18 · Clément Dumas arxiv

Mechanistic interpretability research requires reliable tools for analyzing transformer internals across diverse architectures. Current approaches face a fundamental tradeoff: custom implementations like TransformerLens ensure consistent interfaces but require coding a manual adaptation for each architecture, introducing numerical mismatch with the original models, while direct HuggingFace access through NNsight preserves exact behavior but lacks standardization across models. To bridge this gap, we develop nnterp, a lightweight wrapper around NNsight that provides a unified interface for transformer analysis while preserving original HuggingFace implementations. Through automatic module renaming and comprehensive validation testing, nnterp enables researchers to write intervention code once and deploy it across 50+ model variants spanning 16 architecture families. The library includes built-in implementations of common interpretability methods (logit lens, patchscope, activation steering) and provides direct access to attention probabilities for models that support it. By packaging validation tests with the library, researchers can verify compatibility with custom models locally. nnterp bridges the gap between correctness and usability in mechanistic interpretability tooling.

📄 PDF Abstract BibTeX arXiv:2511.14465

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers

2026-04-15 · Nina Żukowska, Wolfgang Stammer, Bernt Schiele, Jonas Fischer arxiv

Transparency of neural networks' internal reasoning is at the heart of interpretability research, adding to trust, safety, and understanding of these models. The field of mechanistic interpretability has recently focused…

Compact Proofs of Model Performance via Mechanistic Interpretability

2024-06-17 · Jason Gross, Rajashree Agrawal, Thomas Kwa, Euan Ong 외

We propose using mechanistic interpretability -- techniques for reverse engineering model weights into human-interpretable algorithms -- to derive and compactly prove formal guarantees on model performance. We prototype …

model

Interpreting Transformers Through Attention Head Intervention

2026-01-07 · Mason Kadem, Rong Zheng arxiv

Neural networks are growing more capable on their own, but we do not understand their neural mechanisms. Understanding these mechanisms' decision-making processes, or mechanistic interpretability, enables (1) accountabil…

InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques

2024-07-19 · Rohan Gupta, Iván Arcuschin, Thomas Kwa, Adrià Garriga-Alonso

Mechanistic interpretability methods aim to identify the algorithm a neural network implements, but it is difficult to validate such methods when the true algorithm is unknown. This work presents InterpBench, a collectio…

Sparse Autoencoders Can Interpret Randomly Initialized Transformers

2025-01-29 · Thomas Heap, Tim Lawson, Lucy Farnik, Laurence Aitchison

Sparse autoencoders (SAEs) are an increasingly popular technique for interpreting the internal representations of transformers. In this paper, we apply SAEs to 'interpret' random transformers, i.e., transformers where th…