paper-with-me

홈 › Papers

Recoverable but Not Stationary:Local Linear Structures in Weights and Activations

2026-06-09 · Irina Piontkovskaia, Sergey Nikolenko arxiv

Task vectors, LoRA, activation steering, and random search around pretrained weights all suggest that learned behaviour can be controlled by linear directions. We ask which linear structures actually exist and on what scale. In a synthetic multitask transformer and LoRA adapters on DistilGPT-2 / GPT-2 we find strong local low-rank task-gradient structure but reject the fixed-task-plane hypothesis: static bases miss the recovery direction, and the useful basis drifts substantially within 100 steps. However, the first recovery updates form a trajectory-prefix basis capturing 77% of the LoRA recovery displacement. We develop random search theory with a Gaussian local-linear theorem that justifies the effectiveness of random parameter search even in very high dimensions. We also study the relation between parameter perturbations and activation steering: a single gradient step produces an activation shift with 0.58 cosine to a labelled-contrast CAA steering vector, with a similar steering effect on Qwen-0.5B BoolQ statements. We validate our results with experiments on synthetic Transformers and LLMs. Our results suggest that linear structures in trained networks are not global task directions, but evolving local geometries that partially persist across parameter and activation spaces.

📄 PDF Abstract BibTeX arXiv:2606.10929

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Time-Varying Home Field Advantage in Football: Learning from a Non-Stationary Causal Process

2025-06-13 · Minhao Qi, Hengrui Cai, Guanyu Hu, Weining Shen

In sports analytics, home field advantage is a robust phenomenon where the home team wins more games than the away team. However, discovering the causal factors behind home field advantage presents unique challenges due …

Causal DiscoverySports Analytics

Learning Temporally Causal Latent Processes from General Temporal Data

2021-10-11 · Weiran Yao, Yuewen Sun, Alex Ho, Changyin Sun 외

Our goal is to recover time-delayed latent causal variables and identify their relations from measured temporal data. Estimating causally-related latent variables from observations is particularly challenging as the late…

Causal DiscoveryRepresentation LearningVideo Understanding

Linear Speedup in Saddle-Point Escape for Decentralized Non-Convex Optimization

2019-10-30 · Stefan Vlaski, Ali H. Sayed

Under appropriate cooperation protocols and parameter choices, fully decentralized solutions for stochastic optimization have been shown to match the performance of centralized solutions and result in linear speedup (in …

Stochastic Optimization

Collective evolution of weights in wide neural networks

2018-10-09 · Dmitry Yarotsky

We derive a nonlinear integro-differential transport equation describing collective evolution of weights under gradient descent in large-width neural-network-like models. We characterize stationary points of the evolutio…

A Hebbian/Anti-Hebbian Neural Network for Linear Subspace Learning: A Derivation from Multidimensional Scaling of Streaming Data

2015-03-02 · Cengiz Pehlevan, Tao Hu, Dmitri B. Chklovskii

Neural network models of early sensory processing typically reduce the dimensionality of streaming input data. Such networks learn the principal subspace, in the sense of principal component analysis (PCA), by adjusting …