paper-with-me

홈 › Papers

Breaking the Attention Bottleneck

2024-06-16 · Kalle Hilsenbek

Attention-based transformers have become the standard architecture in many deep learning fields, primarily due to their ability to model long-range dependencies and handle variable-length input sequences. However, the attention mechanism with its quadratic complexity is a significant bottleneck in the transformer architecture. This algorithm is only uni-directional in the decoder and converges to a static pattern in over-parametrized decoder-only models. I address this issue by developing a generative function as attention or activation replacement. It still has the auto-regressive character by comparing each token with the previous one. In my test setting with nanoGPT this yields a smaller loss while having a smaller model. The loss further drops by incorporating an average context vector. This concept of attention replacement is distributed under the GNU AGPL v3 license at https://gitlab.com/Bachstelze/causal_generation.

📄 PDF Abstract BibTeX arXiv:2406.10906

Code (1)

https://gitlab.com/bachstelze/causal_generation 공식 구현 pytorch

Tasks

Decoder

Similar Papers 제목 키워드 기반

C2LLM Technical Report: A New Frontier in Code Retrieval via Adaptive Cross-Attention Pooling

2025-12-24 · Jin Qin, Zihan Liao, Ziyin Zhang, Hang Yu 외 arxiv

We present C2LLM - Contrastive Code Large Language Models, a family of code embedding models in both 0.5B and 7B sizes. Building upon Qwen-2.5-Coder backbones, C2LLM adopts a Pooling by Multihead Attention (PMA) module f…

Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting

2019-06-29 · NeurIPS 2019 12 · Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou 외

Time series forecasting is an important problem across many domains, including predictions of solar plant energy output, electricity consumption, and traffic jam situation. In this paper, we propose to tackle such foreca…

Time SeriesTime Series AnalysisTime Series Forecasting

Eigen Attention: Attention in Low-Rank Space for KV Cache Compression

2024-08-10 · Utkarsh Saxena, Gobinda Saha, Sakshi Choudhary, Kaushik Roy

Large language models (LLMs) represent a groundbreaking advancement in the domain of natural language processing due to their impressive reasoning abilities. Recently, there has been considerable interest in increasing t…

Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study

2024-10-23 · Shawn Tan, Songlin Yang, Aaron Courville, Rameswar Panda 외

The self-attention mechanism traditionally relies on the softmax operator, necessitating positional embeddings like RoPE, or position biases to account for token order. But current methods using still face length general…

Sigsoftmax: Reanalysis of the Softmax Bottleneck

2018-05-28 · NeurIPS 2018 12 · Sekitoshi Kanai, Yasuhiro Fujiwara, Yuki Yamanaka, Shuichi Adachi

Softmax is an output activation function for modeling categorical probability distributions in many applications of deep learning. However, a recent study revealed that softmax can be a bottleneck of representational cap…

Language ModelingLanguage Modelling