InAttention: Linear Context Scaling for Transformers
VRAM requirements for transformer models scale quadratically with context length due to the self-attention mechanism. In this paper we modify the decoder-only transformer, replacing self-attention with InAttention, which scales linearly with context length during inference by having tokens attend only to initial states. Benchmarking shows that InAttention significantly reduces VRAM usage during inference, enabling handling of long sequences on consumer GPUs. We corroborate that fine-tuning extends context length efficiently, improving performance on long sequences without high training costs. InAttention offers a scalable solution for long-range dependencies in transformer models, paving the way for further optimization.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingDecoderSimilar Papers 제목 키워드 기반
Context-Scaling versus Task-Scaling in In-Context Learning
Transformers exhibit In-Context Learning (ICL), where these models solve new tasks by using examples in the prompt without additional training. In our work, we identify and analyze two key components of ICL: (1) context-…
In-Context LearningxLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
Scaling laws play a central role in the success of Large Language Models (LLMs), enabling the prediction of model performance relative to compute budgets prior to training. While Transformers have been the dominant archi…
Robust Predictions in Games with Rational Inattention
We derive robust predictions in games involving flexible information acquisition, also known as rational inattention (Sims 2003). These predictions remain accurate regardless of the specific methods players employ to gat…
Linear Transformers are Versatile In-Context Learners
Recent research has demonstrated that transformers, particularly linear attention models, implicitly execute gradient-descent-like algorithms on data provided in-context during their forward inference step. However, thei…
On the solution of the variational optimisation in the rational inattention framework
I analyse the solution method for the variational optimisation problem in the rational inattention framework proposed by Christopher A. Sims. The solution, in general, does not exist, although it may exist in exceptional…