paper-with-me

홈 › Papers

Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling

2024-10-02 · Xiang Hu, Zhihao Teng, Jun Zhao, Wei Wu, Kewei Tu

Despite the success of Transformers, handling long contexts remains challenging due to the limited length generalization and quadratic complexity of self-attention. Thus Transformers often require post-training with a larger attention window, significantly increasing computational and memory costs. In this paper, we propose a novel attention mechanism based on dynamic context, Grouped Cross Attention (GCA), which can generalize to 1000 times the pre-training context length while maintaining the ability to access distant information with a constant attention window size. For a given input sequence, we split it into chunks and use each chunk to retrieve top-k relevant past chunks for subsequent text generation. Specifically, unlike most previous works that use an off-the-shelf retriever, our key innovation allows the retriever to learn how to retrieve past chunks that better minimize the auto-regressive loss of subsequent tokens in an end-to-end manner. Such a mechanism accommodates retrieved chunks with a fixed-size attention window to achieve long-range information access, significantly reducing computational and memory costs during training and inference. Experiments show that GCA-based models achieve near-perfect accuracy in passkey retrieval for 16M context lengths, which is 1000 times the training length.

📄 PDF Abstract BibTeX arXiv:2410.01651

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingRetrievalText Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Length Generalization of Causal Transformers without Position Encoding

2024-04-18 · Jie Wang, Tao Ji, Yuanbin Wu, Hang Yan 외

Generalizing to longer sentences is important for recent Transformer-based language models. Besides algorithms manipulating explicit position features, the success of Transformers without position encodings (NoPE) provid…

Language ModelingLanguage ModellingPositionRetrieval

Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval

2026-05-10 · Zichen Zou, Xiaosong Jia, Zuxuan Wu, Yu-Gang Jiang arxiv

Visual Geometry Grounded Transformer (VGGT) advances 3D reconstruction via scalable Transformer architecture, but the quadratic complexity of global attention prevents long context application. StreamVGGT enables streami…

3D Reconstruction

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

2026-05-26 · Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao, Yiming Du 외 arxiv

Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories, requiring them to locate relevant evidence across interleaved text…

Image RetrievalHead DetectionText Retrieval

Parallel Causal Associative Fields: Gated Sparse Memory for Long-Context Language Modeling

2026-06-09 · Muhammad Ahmed arxiv

Transformers achieve strong language modeling performance by providing direct token-to-token communication paths, but causal self-attention scales quadratically with context length. Recurrent and state-space models reduc…

S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference

2026-01-25 · Qingsen Ma, Dianyun Wang, Yaoye Wang, Lechen Ning 외 arxiv

Large language models are increasingly applied to multi-document and long-form inputs, yet long-context inference remains memory- and noise-inefficient. Key-value (KV) caching scales linearly with context length, while e…