paper-with-me

Papers

Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models

2026-02-04 · Youngji Roh, Hyunjin Cho, Jaehyung Kim arxiv

Large Language Models (LLMs) exhibit highly anisotropic internal representations, often characterized by massive activations, a phenomenon where a small subset of feature dimensions possesses magnitudes significantly larger than the rest. While prior works view these extreme dimensions primarily as artifacts to be managed, we propose a distinct perspective: these dimensions serve as intrinsic interpretable functional units arising from domain specialization. Specifically, we propose a simple magnitude-based criterion to identify Domain-Critical Dimensions in a training-free manner. Our analyses reveal that such dimensions behave as interpretable semantic detectors for symbolic/quantitative patterns or domain-specific terms. In addition, we introduce Critical Dimension Steering, which applies activation steering exclusively to the identified dimensions. Empirical results show that this approach outperforms conventional whole-dimension steering in domain adaptation and jailbreaking scenarios.

📄 PDF Abstract BibTeX arXiv:2603.00029

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Adaptation

Similar Papers 제목 키워드 기반

Direction-Dependent Turning Leads to Anisotropic Diffusion and Persistence

2021-01-12 · Nadia Loy, Thomas Hillen, Kevin John Painter

Cells and organisms follow aligned structures in their environment, a process that can generate persistent migration paths. Kinetic transport equations are a popular modelling tool for describing biological movements at …

Massive Activations in Large Language Models

2024-02-27 · MingJie Sun, Xinlei Chen, J. Zico Kolter, Zhuang Liu

We observe an empirical phenomenon in Large Language Models (LLMs) -- very few activations exhibit significantly larger values than others (e.g., 100,000 times larger). We call them massive activations. First, we demonst…

House of Cards: Massive Weights in LLMs

2024-10-02 · Jaehoon Oh, Seungjun Shin, Dokwan Oh

Massive activations, which manifest in specific feature dimensions of hidden states, introduce a significant bias in large language models (LLMs), leading to an overemphasis on the corresponding token. In this paper, we …

parameter-efficient fine-tuning

Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

2026-05-13 · Evelyn Turri, Davide Bucciarelli, Sara Sarto, Lorenzo Baraldi 외 arxiv

Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood. I…

A Refined Analysis of Massive Activations in LLMs

2025-03-28 · Louis Owen, Nilabhra Roy Chowdhury, Abhay Kumar, Fabian Güra

Motivated in part by their relevance for low-precision training and quantization, massive activations in large language models (LLMs) have recently emerged as a topic of interest. However, existing analyses are limited i…

Quantization