paper-with-me

홈 › Papers

Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention

2024-05-27 · Zhen Qin, Weigao Sun, Dong Li, Xuyang Shen, Weixuan Sun, Yiran Zhong

We present Lightning Attention, the first linear attention implementation that maintains a constant training speed for various sequence lengths under fixed memory consumption. Due to the issue with cumulative summation operations (cumsum), previous linear attention implementations cannot achieve their theoretical advantage in a casual setting. However, this issue can be effectively solved by utilizing different attention calculation strategies to compute the different parts of attention. Specifically, we split the attention calculation into intra-blocks and inter-blocks and use conventional attention computation for intra-blocks and linear attention kernel tricks for inter-blocks. This eliminates the need for cumsum in the linear attention calculation. Furthermore, a tiling technique is adopted through both forward and backward procedures to take full advantage of the GPU hardware. To enhance accuracy while preserving efficacy, we introduce TransNormerLLM (TNL), a new architecture that is tailored to our lightning attention. We conduct rigorous testing on standard and self-collected datasets with varying model sizes and sequence lengths. TNL is notably more efficient than other language models. In addition, benchmark results indicate that TNL performs on par with state-of-the-art LLMs utilizing conventional transformer structures. The source code is released at github.com/OpenNLPLab/TransnormerLLM.

📄 PDF Abstract BibTeX arXiv:2405.17381

Code (1)

opennlplab/transnormerllm 공식 구현 pytorch

Tasks

GPULanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models

2024-01-09 · Zhen Qin, Weigao Sun, Dong Li, Xuyang Shen 외

Linear attention is an efficient attention mechanism that has recently emerged as a promising alternative to conventional softmax attention. With its ability to process tokens in linear computational complexities, linear…

GPU

Taipan: Efficient and Expressive State Space Language Models with Selective Attention

2024-10-24 · Chien Van Nguyen, Huy Huu Nguyen, Thang M. Pham, Ruiyi Zhang 외

Efficient long-context language modeling remains a significant challenge in Natural Language Processing (NLP). While Transformers dominate language tasks, they struggle with long sequences due to quadratic computational …

Computational EfficiencyLanguage ModelingLanguage ModellingMamba+1

Transformer Quality in Linear Time

2022-02-21 · Weizhe Hua, Zihang Dai, Hanxiao Liu, Quoc V. Le

We revisit the design choices in Transformers, and propose methods to address their weaknesses in handling long sequences. First, we propose a simple layer named gated attention unit, which allows the use of a weaker sin…

8kLanguage ModelingLanguage ModellingMasked Language Modeling

ClinicalMamba: A Generative Clinical Language Model on Longitudinal Clinical Notes

2024-03-09 · Zhichao Yang, Avijit Mitra, Sunjae Kwon, Hong Yu

The advancement of natural language processing (NLP) systems in healthcare hinges on language model ability to interpret the intricate information contained within clinical notes. This process often requires integrating …

Few-Shot LearningLanguage ModelingLanguage ModellingMamba

ChordMixer: A Scalable Neural Attention Model for Sequences with Different Lengths

2022-06-12 · Ruslan Khalitov, Tong Yu, Lei Cheng, Zhirong Yang

Sequential data naturally have different lengths in many domains, with some very long sequences. As an important modeling tool, neural attention should capture long-range interaction in such sequences. However, most exis…

ChunkingDocument ClassificationLong-range modelingPosition