paper-with-me

Papers

Provable Length Generalization in Sequence Prediction via Spectral Filtering

2024-11-01 · Annie Marsden, Evan Dogariu, Naman Agarwal, Xinyi Chen, Daniel Suo, Elad Hazan

We consider the problem of length generalization in sequence prediction. We define a new metric of performance in this setting -- the Asymmetric-Regret -- which measures regret against a benchmark predictor with longer context length than available to the learner. We continue by studying this concept through the lens of the spectral filtering algorithm. We present a gradient-based learning algorithm that provably achieves length generalization for linear dynamical systems. We conclude with proof-of-concept experiments which are consistent with our theory.

📄 PDF Abstract BibTeX arXiv:2411.01035

Code (0)

등록된 구현이 없습니다.

Tasks

Prediction

Similar Papers 제목 키워드 기반

On Provable Length and Compositional Generalization

2024-02-07 · Kartik Ahuja, Amin Mansouri

Out-of-distribution generalization capabilities of sequence-to-sequence models can be studied from the lens of two crucial forms of generalization: length generalization -- the ability to generalize to longer sequences t…

DiversityOut-of-Distribution GeneralizationState Space Models

SpectraLDS: Provable Distillation for Linear Dynamical Systems

2025-05-23 · Devan Shah, Shlomo Fortgang, Sofiia Druchyna, Elad Hazan

We present the first provable method for identifying symmetric linear dynamical systems (LDS) with accuracy guarantees that are independent of the systems' state dimension or effective memory. Our approach builds upon re…

Language ModelingLanguage Modelling

Spectral State Space Models

2023-12-11 · Naman Agarwal, Daniel Suo, Xinyi Chen, Elad Hazan

This paper studies sequence modeling for prediction tasks with long range dependencies. We propose a new formulation for state space models (SSMs) based on learning linear dynamical systems with the spectral filtering al…

PredictionState Space Models

On the Provable Generalization of Recurrent Neural Networks

2021-09-29 · NeurIPS 2021 12 · Lifu Wang, Bo Shen, Bo Hu, Xing Cao

Recurrent Neural Network (RNN) is a fundamental structure in deep learning. Recently, some works study the training process of over-parameterized neural networks, and show that over-parameterized networks can learn funct…

Alignment-Sensitive Minimax Rates for Spectral Algorithms with Learned Kernels

2025-09-24 · Dongming Huang, Zhifan Li, Yicheng Li, Qian Lin arxiv

We study spectral algorithms in the setting where kernels are learned from data. We introduce the effective span dimension (ESD), an alignment-sensitive complexity measure that depends jointly on the signal, spectrum, an…