paper-with-me

홈 › Papers

Limitations of Normalization in Attention Mechanism

2025-08-25 · Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State arxiv

This paper investigates the limitations of the normalization in attention mechanisms. We begin with a theoretical framework that enables the identification of the model's selective ability and the geometric separation involved in token selection. Our analysis includes explicit bounds on distances and separation criteria for token vectors under softmax scaling. Through experiments with pre-trained GPT-2 model, we empirically validate our theoretical results and analyze key behaviors of the attention mechanism. Notably, we demonstrate that as the number of selected tokens increases, the model's ability to distinguish informative tokens declines, often converging toward a uniform selection pattern. We also show that gradient sensitivity under softmax normalization presents challenges during training, especially at low temperature settings. These findings advance current understanding of softmax-based attention mechanism and motivate the need for more robust normalization and selection strategies in future attention architectures.

📄 PDF Abstract BibTeX arXiv:2508.17821

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AttenST: A Training-Free Attention-Driven Style Transfer Framework with Pre-Trained Diffusion Models

2025-03-10 · Bo Huang, Wenlun Xu, Qizhuo Han, Haodong Jing 외

While diffusion models have achieved remarkable progress in style transfer tasks, existing methods typically rely on fine-tuning or optimizing pre-trained models during inference, leading to high computational costs and …

Style Transfer

Improving Adversarial Robustness of Attribution via Implicit Regularization

2026-05-28 · Amir Mehrpanah, Matteo Gamba, Hossein Azizpour arxiv

The adversarial robustness of attributions is a fundamental requirement for reliable explainability in deep learning, yet existing approaches typically rely on computationally expensive explicit regularization. In this w…

Adversarial Robustness

NAM: Normalization-based Attention Module

2021-11-24 · NeurIPS Workshop ImageNet_PPF 2021 12 · Yichao Liu, Zongru Shao, Yueyang Teng, Nico Hoffmann

Recognizing less salient features is the key for model compression. However, it has not been investigated in the revolutionary attention mechanisms. In this work, we propose a novel normalization-based attention module (…

Model Compression

DINT Transformer

2025-01-29 · Yueyang Cang, Yuhang Liu, Xiaoteng Zhang, Erlu Zhao 외

DIFF Transformer addresses the issue of irrelevant context interference by introducing a differential attention mechanism that enhances the robustness of local attention. However, it has two critical limitations: the lac…

Information RetrievalLanguage ModelingLanguage Modelling

ELA: Efficient Local Attention for Deep Convolutional Neural Networks

2024-03-02 · Wei Xu, Yi Wan

The attention mechanism has gained significant recognition in the field of computer vision due to its ability to effectively enhance the performance of deep neural networks. However, existing methods often struggle to ef…

Dimensionality Reductionimage-classificationImage Classificationobject-detection+1