paper-with-me

Papers

Engineering Monosemanticity in Toy Models

2022-11-16 · Adam S. Jermyn, Nicholas Schiefer, Evan Hubinger

In some neural networks, individual neurons correspond to natural ``features'' in the input. Such \emph{monosemantic} neurons are of great help in interpretability studies, as they can be cleanly understood. In this work we report preliminary attempts to engineer monosemanticity in toy models. We find that models can be made more monosemantic without increasing the loss by just changing which local minimum the training process finds. More monosemantic loss minima have moderate negative biases, and we are able to use this fact to engineer highly monosemantic models. We are able to mechanistically interpret these models, including the residual polysemantic neurons, and uncover a simple yet surprising algorithm. Finally, we find that providing models with more neurons per layer makes the models more monosemantic, albeit at increased computational cost. These findings point to a number of new questions and avenues for engineering monosemanticity, which we intend to study these in future work.

📄 PDF Abstract BibTeX arXiv:2211.09169

Code (1)

adamjermyn/toy_model_interpretability 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective

2024-06-25 · Hanqi Yan, Yanzheng Xiang, Guangyi Chen, Yifei Wang 외

To better interpret the intrinsic mechanism of large language models (LLMs), recent studies focus on monosemanticity on its basic units. A monosemantic neuron is dedicated to a single and specific concept, which forms a …

DiversityFeature Correlation

Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness

2024-10-27 · Qi Zhang, Yifei Wang, Jingyi Cui, Xiang Pan 외

Deep learning models often suffer from a lack of interpretability due to polysemanticity, where individual neurons are activated by multiple unrelated semantics, resulting in unclear attributions of model behavior. Recen…

Domain GeneralizationFew-Shot Learning

Learning from Emergence: A Study on Proactively Inhibiting the Monosemantic Neurons of Artificial Neural Networks

2023-12-17 · Jiachuan Wang, Shimin Di, Lei Chen, Charles Wang Wai Ng

Recently, emergence has received widespread attention from the research community along with the success of large-scale models. Different from the literature, we hypothesize a key factor that promotes the performance dur…

Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models

2025-04-03 · Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge Belongie 외

Sparse Autoencoders (SAEs) have recently been shown to enhance interpretability and steerability in Large Language Models (LLMs). In this work, we extend the application of SAEs to Vision-Language Models (VLMs), such as …

Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims

2026-05-08 · Zezheng Lin, Fengming Liu arxiv

Mechanistic interpretability papers increasingly use causal vocabulary: circuits, mediators, causal abstraction, monosemanticity. Such claims require explicit identification assumptions. A purposive audit of 10 papers ac…