paper-with-me

홈 › Papers

Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure

2026-05-09 · Nilesh Sarkar, Dawar Jyoti Deka arxiv

Sparse autoencoders (SAEs) decompose transformer residual streams into interpretable feature dictionaries, yet the relationship between SAE width and causal influence on model output has not been systematically characterised. We introduce causal dimensionality kappa(L, M, T), defined as the effective rank of the expected Jacobian outer product at layer L, and show it can be estimated via the SAE width sweep paired with attribution patching. Across seven SAE widths from 16,384 to 1,048,576 features on Gemma-2-2B layer 12, representational capacity grows 15.6x while causal capacity grows only 4.35x: a robust separation we term the representational-causal wedge. A saturating fit yields kappa-hat approximately 1,990 with kappa-hat / d_model = 0.86 and participation-ratio lower bound kappa_PR approximately 280. Crucially, kappa is invariant to model scaling: Gemma-2-9B and Gemma-2-2B yield identical N_causal = 328 at the same SAE width despite a 3.46x parameter increase (the count is forced to 2% of SAE width by calibration; the substantive empirical claim is shape invariance of the AtP score distribution under matched seq=512 conditions). Across eight network depths kappa is constant while the absolute attribution threshold drops 20x from layer 1 to layer 23. Five controls (architecture invariance, threshold robustness, geometric privilege, synthetic ground-truth recovery, and a four-cell encoder/decoder ablation) pin down what kappa measures and what it does not. Our findings establish kappa as a measurable, model-intrinsic property of transformer layers: sub-linearly recoverable by SAE width, invariant to model scaling, and structured across network depth.

📄 PDF Abstract BibTeX arXiv:2605.08740

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Transformer Is Inherently a Causal Learner

2026-01-09 · Xinyue Wang, Stephen Wang, Biwei Huang arxiv

We reveal that transformers trained in an autoregressive manner naturally encode time-delayed causal structures in their learned representations. When predicting future values in multivariate time series, the gradient se…

Multi-resolution Enhancement for Full Spectrum Neural Representations

2025-09-19 · Yuan Ni, Zhantao Chen, Shizhou Xu, Cheng Peng 외 arxiv

Scientific data acquisition continues to outpace storage and analysis capabilities, making voxel-based representations increasingly intractable. Implicit neural representations (INRs) offer a promising solution by encodi…

Identifying Weight-Variant Latent Causal Models

2022-08-30 · Yuhang Liu, Zhen Zhang, Dong Gong, Mingming Gong 외

The task of causal representation learning aims to uncover latent higher-level causal representations that affect lower-level observations. Identifying true latent causal representations from observed data, while allowin…

Representation Learning

Collaborative causal inference on distributed data

2022-08-16 · Yuji Kawamata, Ryoki Motai, Yukihiko Okada, Akira Imakura 외

In recent years, the development of technologies for causal inference with privacy preservation of distributed data has gained considerable attention. Many existing methods for distributed data focus on resolving the lac…

Causal InferenceDimensionality Reduction

When Dimensionality Hurts: The Role of LLM Embedding Compression for Noisy Regression Tasks

2025-02-04 · Felix Drinkall, Janet B. Pierrehumbert, Stefan Zohren

Large language models (LLMs) have shown remarkable success in language modelling due to scaling laws found in model size and the hidden dimension of the model's text representation. Yet, we demonstrate that compressed re…

Language Modelling