paper-with-me

홈 › Papers

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

2026-06-24 · Praneet Suresh, Jack Stanley, Sonia Joseph, Luca Scimeca, Danilo Bzdok arxiv

Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-world deployments often face unexpected or adversarial data that diverges from training data distributions. Without explicit mechanisms for handling such shifts, model reliability and safety degrade, urging more disciplined study of out-of-distribution (OOD) settings for transformers. By systematic experiments, we present a mechanistic framework for delineating the precise contours of transformer model robustness. We find that OOD inputs, including subtle typos and jailbreak prompts, drive language models to operate on an increased number of fallacious concepts in their internals. We leverage this device to quantify and understand the degree of distributional shift in prompts, enabling a mechanistically grounded fine-tuning strategy to robustify LLMs. Expanding the very notion of OOD from input data to a model's private computational processes, a new transformer diagnostic at inference time is a critical step toward making AI systems safe for deployment across science, business, and government.

📄 PDF Abstract BibTeX arXiv:2606.26396

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders

2024-11-20 · Charles O'Neill, Alim Gumran, David Klindt

A recent line of work has shown promise in using sparse autoencoders (SAEs) to uncover interpretable features in neural network representations. However, the simple linear-nonlinear encoding mechanism in SAEs limits thei…

compressed sensingLanguage ModelingLanguage ModellingLarge Language Model

How LLMs Learn: Tracing Internal Representations with Sparse Autoencoders

2025-03-09 · Tatsuro Inaba, Kentaro Inui, Yusuke Miyao, Yohei Oseki 외

Large Language Models (LLMs) demonstrate remarkable multilingual capabilities and broad knowledge. However, the internal mechanisms underlying the development of these capabilities remain poorly understood. To investigat…

Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit

2025-06-05 · Valérie Costa, Thomas Fel, Ekdeep Singh Lubana, Bahareh Tolooshams 외

Sparse autoencoders (SAEs) have recently become central tools for interpretability, leveraging dictionary learning principles to extract sparse, interpretable features from neural representations whose underlying structu…

Dictionary Learning

GPT and Prejudice: A Sparse Approach to Understanding Learned Representations in Large Language Models

2025-09-24 · Mariam Mahran, Katharina Simbeck arxiv

Large Language Models (LLMs) are trained on massive, unstructured corpora, making it unclear which social patterns and biases they absorb and later reproduce. Existing evaluations typically examine outputs or activations…

SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders

2024-10-09 · Constantin Venhoff, Anisoara Calinescu, Philip Torr, Christian Schroeder de Witt

A key challenge in interpretability is to decompose model activations into meaningful features. Sparse autoencoders (SAEs) have emerged as a promising tool for this task. However, a central problem in evaluating the qual…