paper-with-me

Papers

Evaluating SAE interpretability without explanations

2025-07-11 · Gonçalo Paulo, Nora Belrose arxiv

Sparse autoencoders (SAEs) and transcoders have become important tools for machine learning interpretability. However, measuring how interpretable they are remains challenging, with weak consensus about which benchmarks to use. Most evaluation procedures start by producing a single-sentence explanation for each latent. These explanations are then evaluated based on how well they enable an LLM to predict the activation of a latent in new contexts. This method makes it difficult to disentangle the explanation generation and evaluation process from the actual interpretability of the latents discovered. In this work, we adapt existing methods to assess the interpretability of sparse coders, with the advantage that they do not require generating natural language explanations as an intermediate step. This enables a more direct and potentially standardized assessment of interpretability. Furthermore, we compare the scores produced by our interpretability metrics with human evaluations across similar tasks and varying setups, offering suggestions for the community on improving the evaluation of these techniques.

📄 PDF Abstract BibTeX arXiv:2507.08473

Code (0)

등록된 구현이 없습니다.

Tasks

Explanation Generation

Similar Papers 제목 키워드 기반

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

2025-05-02 · Kola Ayonrinde, Louis Jaburi

Mechanistic Interpretability (MI) aims to understand neural networks through causal explanations. Though MI has many explanation-generating methods, progress has been limited by the lack of a universal approach to evalua…

Philosophy

QUACKIE: A NLP Classification Task With Ground Truth Explanations

2020-12-24 · Yves Rychener, Xavier Renard, Djamé Seddah, Pascal Frossard 외

NLP Interpretability aims to increase trust in model predictions. This makes evaluating interpretability approaches a pressing issue. There are multiple datasets for evaluating NLP Interpretability, but their dependence …

ClassificationGeneral ClassificationQuestion Answering

HIVE: Evaluating the Human Interpretability of Visual Explanations

2021-12-06 · Sunnie S. Y. Kim, Nicole Meister, Vikram V. Ramaswamy, Ruth Fong 외

As AI technology is increasingly applied to high-impact, high-risk domains, there have been a number of new methods aimed at making AI models more human interpretable. Despite the recent growth of interpretability work, …

Decision MakingDiversity

Evaluating the Robustness of Interpretability Methods through Explanation Invariance and Equivariance

2023-04-13 · NeurIPS 2023 11 · Jonathan Crabbé, Mihaela van der Schaar

Interpretability methods are valuable only if their explanations faithfully describe the explained model. In this work, we consider neural networks whose predictions are invariant under a specific symmetry group. This in…

Pitfalls in Evaluating Interpretability Agents

2026-03-20 · Tal Haklay, Nikhil Prakash, Sana Pandey, Antonio Torralba 외 arxiv

Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverage large language models (LLMs) at increa…