paper-with-me

홈 › Papers

On the Pitfalls of Analyzing Individual Neurons in Language Models

2021-10-14 · ICLR 2022 4 · Omer Antverg, Yonatan Belinkov

While many studies have shown that linguistic information is encoded in hidden word representations, few have studied individual neurons, to show how and in which neurons it is encoded. Among these, the common approach is to use an external probe to rank neurons according to their relevance to some linguistic attribute, and to evaluate the obtained ranking using the same probe that produced it. We show two pitfalls in this methodology: 1. It confounds distinct factors: probe quality and ranking quality. We separate them and draw conclusions on each. 2. It focuses on encoded information, rather than information that is used by the model. We show that these are not the same. We compare two recent ranking methods and a simple one we introduce, and evaluate them with regard to both of these aspects.

📄 PDF Abstract BibTeX arXiv:2110.07483

Code (2)

omerant/individual-neurons 공식 구현 pytorch
technion-cs-nlp/individual-neurons-pitfalls 공식 구현 pytorch

Tasks

Attribute

Similar Papers 제목 키워드 기반

Building population models for large-scale neural recordings: opportunities and pitfalls

2021-02-03 · Cole Hurwitz, Nina Kudryashova, Arno Onken, Matthias H. Hennig

Modern recording technologies now enable simultaneous recording from large numbers of neurons. This has driven the development of new statistical models for analyzing and interpreting neural population activity. Here we …

Analyzing Individual Neurons in Pre-trained Language Models

2020-10-06 · EMNLP 2020 11 · Nadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan Belinkov

While a lot of analysis has been carried to demonstrate linguistic knowledge captured by the representations learned within deep NLP models, very little attention has been paid towards individual neurons.We carry outa ne…

What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models

2018-12-21 · Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov 외

Despite the remarkable evolution of deep neural networks in natural language processing (NLP), their interpretability remains a challenge. Previous work largely focused on what these models learn at the representation le…

Language ModelingLanguage ModellingMachine TranslationNMT+1

LM Transparency Tool: Interactive Tool for Analyzing Transformer Language Models

2024-04-10 · Igor Tufanov, Karen Hambardzumyan, Javier Ferrando, Elena Voita

We present the LM Transparency Tool (LM-TT), an open-source interactive toolkit for analyzing the internal workings of Transformer-based language models. Differently from previously existing tools that focus on isolated …

Decision Making

NeuroX: A Toolkit for Analyzing Individual Neurons in Neural Networks

2018-12-21 · Fahim Dalvi, Avery Nortonsmith, D. Anthony Bau, Yonatan Belinkov 외

We present a toolkit to facilitate the interpretation and understanding of neural network models. The toolkit provides several methods to identify salient neurons with respect to the model itself or an external task. A u…