paper-with-me

홈 › Papers

Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences

2026-04-22 · Neehal Tumma, Noel Loo, Daniela Rus arxiv

To address the increasing long-context compute limitations of softmax attention, several subquadratic recurrent operators have been developed. This work includes models such as Mamba-2, DeltaNet, Gated DeltaNet (GDN), and Kimi Delta Attention (KDA). As the space of recurrences grows, a parallel line of work has arisen to taxonomize them. One compelling view is the test-time regression (TTR) framework, which interprets recurrences as performing online least squares updates that learn a linear map from the keys to values. Existing delta-rule recurrences can be seen as first-order approximations to this objective, but notably ignore the curvature of the least-squares loss during optimization. In this work, we address this by introducing preconditioning to these recurrences. Starting from the theory of online least squares, we derive equivalences between linear attention and the delta rule in the exactly preconditioned case. Next, we realize this theory in practice by proposing a diagonal approximation: this enables us to introduce preconditioned variants of DeltaNet, GDN, and KDA alongside efficient chunkwise parallel algorithms for computing them. Empirically, we find that our preconditioned delta-rule recurrences yield consistent performance improvements across synthetic recall benchmarks and language modeling at the 340M and 1B scale.

📄 PDF Abstract BibTeX arXiv:2604.21100

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Taming Curvature: Architecture Warm-Up for Stable Transformer Training

2026-06-15 · Sameera Ramasinghe, Ajanthan Thalaiyasingam, Hadi Mohaghegh Dolatabadi, Chamin Hewa Koneputugodage 외 arxiv

Training billion-parameter Transformers is often brittle, with transient loss spikes and divergence that waste compute. Even though the recently developed Edge of Stability (EoS) theory provides a powerful tool to unders…

Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

2026-05-21 · Ali Hatamizadeh, Yejin Choi, Jan Kautz arxiv

Linear attention replaces the unbounded cache of softmax attention with a fixed-size recurrent state, reducing sequence mixing to linear time and decoding to constant memory. The hard part is not just what to forget, but…

PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective

2025-05-27 · Tim Tsz-Kit Lau, Qi Long, Weijie Su

The ever-growing scale of deep learning models and datasets underscores the critical importance of efficient optimization methods. While preconditioned gradient methods such as Adam and AdamW are the de facto optimizers …

Language ModelingLanguage Modelling

Adaptively Preconditioned Stochastic Gradient Langevin Dynamics

2019-06-10 · Chandrasekaran Anirudh Bhardwaj

Stochastic Gradient Langevin Dynamics infuses isotropic gradient noise to SGD to help navigate pathological curvature in the loss landscape for deep networks. Isotropic nature of the noise leads to poor scaling, and adap…

Navigate

Design Criteria for SGD Preconditioners: Local Conditioning, Noise Floors, and Basin Stability

2025-11-24 · Mitchell Scott, Tianshi Xu, Ziyuan Tang, Alexandra Pichette-Emmons 외 arxiv

Stochastic Gradient Descent (SGD) often slows in the late stage of training due to anisotropic curvature and gradient noise. We analyze preconditioned SGD in the geometry induced by a symmetric positive definite matrix $…