paper-with-me

Papers

Analyzing (In)Abilities of SAEs via Formal Languages

2024-10-15 · Abhinav Menon, Manish Shrivastava, David Krueger, Ekdeep Singh Lubana

Autoencoders have been used for finding interpretable and disentangled features underlying neural network representations in both image and text domains. While the efficacy and pitfalls of such methods are well-studied in vision, there is a lack of corresponding results, both qualitative and quantitative, for the text domain. We aim to address this gap by training sparse autoencoders (SAEs) on a synthetic testbed of formal languages. Specifically, we train SAEs on the hidden representations of models trained on formal languages (Dyck-2, Expr, and English PCFG) under a wide variety of hyperparameter settings, finding interpretable latents often emerge in the features learned by our SAEs. However, similar to vision, we find performance turns out to be highly sensitive to inductive biases of the training pipeline. Moreover, we show latents correlating to certain features of the input do not always induce a causal impact on model's computation. We thus argue that causality has to become a central target in SAE training: learning of causal features should be incentivized from the ground-up. Motivated by this, we propose and perform preliminary investigations for an approach that promotes learning of causally relevant features in our formal language setting.

📄 PDF Abstract BibTeX arXiv:2410.11767

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Universal Vibe? Finding and Controlling Language-Agnostic Informal Register with SAEs

2026-03-27 · Uri Z. Kialy, Avi Shtarkberg, Ayal Klein arxiv

While multilingual language models successfully transfer factual and syntactic knowledge across languages, it remains unclear whether they process culture-specific pragmatic registers, such as slang, as isolated language…

Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders

2025-05-08 · Boyi Deng, Yu Wan, Yidan Zhang, Baosong Yang 외

The mechanisms behind multilingual capabilities in Large Language Models (LLMs) have been examined using neuron-based or internal-activation-based methods. However, these methods often face challenges such as superpositi…

Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection

2026-04-29 · Ahyoung Oh, Wonseok Shin, Songkuk Kim arxiv

Sparse Autoencoders (SAEs) have demonstrated significant success in interpreting Large Language Models (LLMs) by decomposing dense representations into sparse, semantic components. However, their potential for analyzing …

Out-of-Distribution Detection

TopK Language Models

2025-06-26 · Ryosuke Takahashi, Tatsuro Inaba, Kentaro Inui, Benjamin Heinzerling

Sparse autoencoders (SAEs) have become an important tool for analyzing and interpreting the activation space of transformer-based language models (LMs). However, SAEs suffer several shortcomings that diminish their utili…

Computational Efficiency

In What Languages are Generative Language Models the Most Formal? Analyzing Formality Distribution across Languages

2023-02-23 · Asım Ersoy, Gerson Vizcarra, Tasmiah Tahsin Mayeesha, Benjamin Muller

Multilingual generative language models (LMs) are increasingly fluent in a large variety of languages. Trained on the concatenation of corpora in multiple languages, they enable powerful transfer from high-resource langu…

Cultural Vocal Bursts Intensity PredictionDiversity