paper-with-me

홈 › Papers

Superiority of Softmax: Unveiling the Performance Edge Over Linear Attention

2023-10-18 · Yichuan Deng, Zhao Song, Tianyi Zhou

Large transformer models have achieved state-of-the-art results in numerous natural language processing tasks. Among the pivotal components of the transformer architecture, the attention mechanism plays a crucial role in capturing token interactions within sequences through the utilization of softmax function. Conversely, linear attention presents a more computationally efficient alternative by approximating the softmax operation with linear complexity. However, it exhibits substantial performance degradation when compared to the traditional softmax attention mechanism. In this paper, we bridge the gap in our theoretical understanding of the reasons behind the practical performance gap between softmax and linear attention. By conducting a comprehensive comparative analysis of these two attention mechanisms, we shed light on the underlying reasons for why softmax attention outperforms linear attention in most scenarios.

📄 PDF Abstract BibTeX arXiv:2310.11685

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

AdaDistill: Adaptive Knowledge Distillation for Deep Face Recognition

2024-07-01 · Fadi Boutros, Vitomir Štruc, Naser Damer

Knowledge distillation (KD) aims at improving the performance of a compact student model by distilling the knowledge from a high-performing teacher model. In this paper, we present an adaptive KD approach, namely AdaDist…

Face RecognitionKnowledge Distillation

Ensemble Soft-Margin Softmax Loss for Image Classification

2018-05-10 · Xiaobo Wang, Shifeng Zhang, Zhen Lei, Si Liu 외

Softmax loss is arguably one of the most popular losses to train CNN models for image classification. However, recent works have exposed its limitation on feature discriminability. This paper casts a new viewpoint on the…

ClassificationDiversityGeneral Classificationimage-classification+1

On Softmax Direct Preference Optimization for Recommendation

2024-06-13 · Yuxin Chen, Junfei Tan, An Zhang, Zhengyi Yang 외

Recommender systems aim to predict personalized rankings based on user preference data. With the rise of Language Models (LMs), LM-based recommenders have been widely explored due to their extensive world knowledge and p…

Language ModelingLanguage ModellingRecommendation SystemsWorld Knowledge

Noisy Softmax: Improving the Generalization Ability of DCNN via Postponing the Early Softmax Saturation

2017-08-12 · CVPR 2017 7 · Binghui Chen, Weihong Deng, Junping Du

Over the past few years, softmax and SGD have become a commonly used component and the default training strategy in CNN frameworks, respectively. However, when optimizing CNNs with SGD, the saturation behavior behind sof…

$ε$-Softmax: Approximating One-Hot Vectors for Mitigating Label Noise

2025-08-04 · Jialiang Wang, Xiong Zhou, Deming Zhai, Junjun Jiang 외 arxiv

Noisy labels pose a common challenge for training accurate deep neural networks. To mitigate label noise, prior studies have proposed various robust loss functions to achieve noise tolerance in the presence of label nois…