paper-with-me

홈 › Papers

Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond

2024-10-09 · Dilyara Bareeva, Galip Ümit Yolcu, Anna Hedström, Niklas Schmolenski, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin

In recent years, training data attribution (TDA) methods have emerged as a promising direction for the interpretability of neural networks. While research around TDA is thriving, limited effort has been dedicated to the evaluation of attributions. Similar to the development of evaluation metrics for traditional feature attribution approaches, several standalone metrics have been proposed to evaluate the quality of TDA methods across various contexts. However, the lack of a unified framework that allows for systematic comparison limits trust in TDA methods and stunts their widespread adoption. To address this research gap, we introduce Quanda, a Python toolkit designed to facilitate the evaluation of TDA methods. Beyond offering a comprehensive set of evaluation metrics, Quanda provides a uniform interface for seamless integration with existing TDA implementations across different repositories, thus enabling systematic benchmarking. The toolkit is user-friendly, thoroughly tested, well-documented, and available as an open-source library on PyPi and under https://github.com/dilyabareeva/quanda.

📄 PDF Abstract BibTeX arXiv:2410.07158

Code (1)

dilyabareeva/quanda 공식 구현 pytorch

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Library 설명 없음

Similar Papers 제목 키워드 기반

reward-lens: A Mechanistic Interpretability Library for Reward Models

2026-04-28 · Mohammed Suhail B Nadaf arxiv

Every RLHF-trained language model is shaped by a reward model, yet the mechanistic interpretability toolkit -- logit lens, direct logit attribution, activation patching, sparse autoencoders -- was built for generative LL…

Towards Answering Open-ended Ethical Quandary Questions

2022-05-12 · Yejin Bang, Nayeon Lee, Tiezheng Yu, Leila Khalatbari 외

Considerable advancements have been made in various NLP tasks based on the impressive power of large language models (LLMs) and many NLP applications are deployed in our daily lives. In this work, we challenge the capabi…

Few-Shot LearningGenerative Question AnsweringNatural Language UnderstandingPhilosophy+1

Inseq: An Interpretability Toolkit for Sequence Generation Models

2023-02-27 · Gabriele Sarti, Nils Feldhus, Ludwig Sickert, Oskar van der Wal 외

Past work in natural language processing interpretability focused mainly on popular classification tasks while largely overlooking generation settings, partly due to a lack of dedicated tools. In this work, we introduce …

DecoderFeature ImportanceMachine TranslationText Generation+1

The Dead Salmons of AI Interpretability

2025-12-21 · Maxime Méloux, Giada Dirupo, François Portet, Maxime Peyrard arxiv

In a striking neuroscience study, the authors placed a dead salmon in an MRI scanner and showed it images of humans in social situations. Astonishingly, standard analyses of the time reported brain regions predictive of …

Neutrally Evolving Interlocking Complexity in the Quandary Den

2026-04-20 · Andrew Walsh arxiv

Molecular biology features numerous complexes of proteins that coordinate in an interlocking fashion to fulfill different functions. Adaptive evolution explains some of this complexity, but needn't be the default when ne…