paper-with-me

Papers

Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers

2024-10-08 · Sanchayan Dutta, Suvrit Sra

We show that memory-augmented Transformers can implement the entire class of linear first-order methods (LFOMs), a class that contains gradient descent (GD) and more advanced methods such as conjugate gradient descent (CGD), momentum methods and all other variants that linearly combine past gradients. Building on prior work that studies how Transformers simulate GD, we provide theoretical and empirical evidence that memory-augmented Transformers can learn more advanced algorithms. We then take a first step toward turning the learned algorithms into actually usable methods by developing a mixture-of-experts (MoE) approach for test-time adaptation to out-of-distribution (OOD) samples. Lastly, we show that LFOMs can themselves be treated as learnable algorithms, whose parameters can be learned from data to attain strong performance.

📄 PDF Abstract BibTeX arXiv:2410.07263

Code (0)

등록된 구현이 없습니다.

Tasks

AllMixture-of-ExpertsTest-time Adaptation

Methods 이 논문이 사용한 방법론

Linear Regression Linear Regression is a method for modelling a relationship between a dependent variable and independent variables. These models can be fit with numerous approaches. The most…

Similar Papers 제목 키워드 기반

Efficient Convex Optimization Requires Superlinear Memory

2022-03-29 · Annie Marsden, Vatsal Sharan, Aaron Sidford, Gregory Valiant

We show that any memory-constrained, first-order algorithm which minimizes $d$-dimensional, $1$-Lipschitz convex functions over the unit ball to $1/\mathrm{poly}(d)$ accuracy using at most $d^{1.25 - \delta}$ bits of mem…

Memory-Constrained No-Regret Learning in Adversarial Bandits

2020-02-26 · Xiao Xu, Qing Zhao

An adversarial bandit problem with memory constraints is studied where only the statistics of a subset of arms can be stored. A hierarchical learning policy that requires only a sublinear order of memory space in terms o…

S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring

2026-07-30 · Glenn Anta Bucagu, Thorir Mar Ingolfsson, Yawei Li, Luca Benini arxiv

Foundation models offer a promising paradigm for Electroencephalography (EEG) analysis, leveraging generalizable representations from vast unlabeled datasets. Yet, Transformer-based architectures face a critical bottlene…

Enhancing Transformers for Generalizable First-Order Logical Entailment

2025-01-01 · Tianshi Zheng, Jiazheng Wang, ZiHao Wang, Jiaxin Bai 외

Transformers, as a fundamental deep learning architecture, have demonstrated remarkable capabilities in reasoning. This paper investigates the generalizable first-order logical reasoning ability of transformers with thei…

Logical ReasoningOut-of-Distribution Generalization

Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics

2026-07-05 · Jakob Kienegger, Tal Peer, Sina Khanagha, Timo Gerkmann arxiv

Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization. Modern beamformers are often parameterized by deep neural n…

Speech Enhancement