paper-with-me

홈 › Papers

Structural Instability of Feature Composition

2026-04-18 · Yunpeng Zhou arxiv

Sparse Autoencoders (SAEs) have emerged as a powerful paradigm for disentangling feature superposition in transformer-based architectures, enabling precise control via activation steering. However, the theoretical foundations of compositional steering -- the simultaneous activation of distinct semantic latents -- remain under-explored. The prevailing Linear Representation Hypothesis often abstracts away non-linear interference effects that arise in overcomplete dictionaries. We present a geometric framework for analyzing the instability of feature unions. Modeling the activation space as a high-dimensional sparse cone manifold, we derive an asymptotic compositional-collapse threshold under a spherical dictionary model, characterized by the Gaussian mean width (statistical dimension) of the signal cone. We further show that, in the high-bias regime, ReLU rectification converts microscopic correlation-induced variance fluctuations into a systematic drift that accumulates under composition, yielding interference growth consistent with a ratchet effect. We validate the predicted scaling trends on structured semantic features extracted from CLEVR, where hierarchical correlations accelerate the transition relative to random baselines. Together, our results highlight geometric constraints on the scalability of union-based steering and motivate composition mechanisms that explicitly manage interference beyond naive linear superposition.

📄 PDF Abstract BibTeX arXiv:2605.05223

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Donor-Aware scRNA-seq Benchmarks for IBD Classification

2026-05-05 · Jonathan Muhire arxiv

Donor-level disease classification from single-cell RNA sequencing (scRNA-seq) requires strict donor-aware cross-validation: naive pipelines that split cells randomly conflate training and test donors, inflating reported…

Frequency-Enhanced Dual-Subspace Networks for Few-Shot Fine-Grained Image Classification

2026-04-16 · Meijia Wang, Guochao Wang, Haozhen Chu, Bin Yao 외 arxiv

Few-shot fine-grained image classification aims to recognize subcategories with high visual similarity using only a limited number of annotated samples. Existing metric learning-based methods typically rely solely on spa…

Fine-Grained Image ClassificationFine-Grained Visual RecognitionComputational EfficiencyMetric Learning

Causal Schrödinger Bridges: Constrained Optimal Transport on Structural Manifolds

2026-02-09 · Rui Wu, Li YongJun arxiv

Generative modeling typically seeks the path of least action via deterministic flows (ODE). While effective for in-distribution tasks, we argue that these deterministic paths become brittle under causal interventions, wh…

Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation

2025-09-20 · Zuhair Hasan Shaik, Abdullah Mazhar, Aseem Srivastava, Md Shad Akhtar arxiv

Large Language Models have demonstrated impressive fluency across diverse tasks, yet their tendency to produce toxic content remains a critical challenge for AI safety and public trust. Existing toxicity mitigation appro…

Geometric Decoupling: Diagnosing the Structural Instability of Latent

2026-04-20 · Yuanbang Liang, Zhengwen Chen, Yu-Kun Lai arxiv

Latent Diffusion Models (LDMs) achieve high-fidelity synthesis but suffer from latent space brittleness, causing discontinuous semantic jumps during editing. We introduce a Riemannian framework to diagnose this instabili…