paper-with-me

홈 › Papers

Grouped self-attention mechanism for a memory-efficient Transformer

2022-10-02 · Bumjun Jung, Yusuke Mukuta, Tatsuya Harada

Time-series data analysis is important because numerous real-world tasks such as forecasting weather, electricity consumption, and stock market involve predicting data that vary over time. Time-series data are generally recorded over a long period of observation with long sequences owing to their periodic characteristics and long-range dependencies over time. Thus, capturing long-range dependency is an important factor in time-series data forecasting. To solve these problems, we proposed two novel modules, Grouped Self-Attention (GSA) and Compressed Cross-Attention (CCA). With both modules, we achieved a computational space and time complexity of order $O(l)$ with a sequence length $l$ under small hyperparameter limitations, and can capture locality while considering global information. The results of experiments conducted on time-series datasets show that our proposed model efficiently exhibited reduced computational complexity and performance comparable to or better than existing methods.

📄 PDF Abstract BibTeX arXiv:2210.00440

Code (0)

등록된 구현이 없습니다.

Tasks

Time SeriesTime Series Analysis

Similar Papers 제목 키워드 기반

Weighted Grouped Query Attention in Transformers

2024-07-15 · Sai Sena Chinnakonduru, Astarag Mohapatra

The attention mechanism forms the foundational blocks for transformer language models. Recent approaches show that scaling the model achieves human-level performance. However, with increasing demands for scaling and cons…

Decoder

Beyond Uniform Query Distribution: Key-Driven Grouped Query Attention

2024-08-15 · Zohaib Khan, Muhammad Khaquan, Omer Tafveez, Burhanuddin Samiwala 외

The Transformer architecture has revolutionized deep learning through its Self-Attention mechanism, which effectively captures contextual information. However, the memory footprint of Self-Attention presents significant …

image-classificationImage Classification

Is Attention All What You Need? -- An Empirical Investigation on Convolution-Based Active Memory and Self-Attention

2019-12-27 · Thomas Dowdell, Hongyu Zhang

The key to a Transformer model is the self-attention mechanism, which allows the model to analyze an entire sequence in a computationally efficient manner. Recent work has suggested the possibility that general attention…

AllLanguage Modelling

Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling

2024-10-02 · Xiang Hu, Zhihao Teng, Jun Zhao, Wei Wu 외

Despite the success of Transformers, handling long contexts remains challenging due to the limited length generalization and quadratic complexity of self-attention. Thus Transformers often require post-training with a la…

Language ModelingLanguage ModellingRetrievalText Generation

Green Hierarchical Vision Transformer for Masked Image Modeling

2022-05-26 · Lang Huang, Shan You, Mingkai Zheng, Fei Wang 외

We present an efficient approach for Masked Image Modeling (MIM) with hierarchical Vision Transformers (ViTs), allowing the hierarchical ViTs to discard masked patches and operate only on the visible ones. Our approach c…

GPUObject Detection