paper-with-me

Papers

Taming the Exponential: A Fast Softmax Surrogate for Integer-Native Edge Inference

2026-04-02 · Dimitrios Danopoulos, Enrico Lupi, Michael Kagan, Maurizio Pierini arxiv

Softmax can become a computational bottleneck in the Transformer model's Multi-Head Attention (MHA) block, particularly in small models under low-precision inference, where exponentiation and normalization incur significant overhead. As such, we suggest using Head-Calibrated Clipped-Linear Softmax (HCCS), a bounded, monotone surrogate to the exponential softmax function, which uses a clipped linear mapping of the max centered attention logits. This approximation produces a stable probability distribution, maintains the ordering of the original logits and has non-negative values. HCCS differs from previous softmax surrogates as it includes a set of lightweight calibration parameters that are optimized offline based on a representative dataset and calibrated for each individual attention head to preserve the statistical properties of the individual heads. We describe a hardware-motivated implementation of HCCS for high-throughput scenarios targeting the AMD Versal AI Engines. The current reference implementations from AMD for this platform rely upon either bfloat16 arithmetic or LUTs to perform the exponential operation, which might limit the throughput of the platform and fail to utilize the high-throughput integer vector processing units of the AI Engine. In contrast, HCCS provides a natural mapping to the AI Engines' int8 multiply accumulate (MAC) units. To the best of our knowledge, this is the first int8 optimized softmax surrogate for AMD AI engines that significantly exceeds the speed performance of other reference implementations while maintaining competitive task accuracy on small or heavily quantized MHA workloads after quantization-aware retraining.

📄 PDF Abstract BibTeX arXiv:2604.02292

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference

2025-11-26 · Wanli Zhong, Haibo Feng, Zirui Zhou, Hanyang Peng 외 arxiv

Deploying Transformer models on edge devices is limited by latency and energy budgets. While INT8 quantization effectively accelerates the primary matrix multiplications, it exposes the softmax-related path as the domina…

QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention

2026-04-28 · Sehyeon Oh, Yongin Kwon, Jemin Lee arxiv

FlashAttention improves efficiency through tiling, but its online softmax still relies on floating-point arithmetic for numerical stability, making full quantization difficult. We identify three main obstacles to integer…

PSL: Rethinking and Improving Softmax Loss from Pairwise Perspective for Recommendation

2024-10-31 · Weiqin Yang, Jiawei Chen, Xin Xin, Sheng Zhou 외

Softmax Loss (SL) is widely applied in recommender systems (RS) and has demonstrated effectiveness. This work analyzes SL from a pairwise perspective, revealing two significant limitations: 1) the relationship between SL…

Recommendation Systems

Taming the Exponential Action Set: Sublinear Regret and Fast Convergence to Nash Equilibrium in Online Congestion Games

2023-06-19 · Jing Dong, Jingyu Wu, Siwei Wang, Baoxiang Wang 외

The congestion game is a powerful model that encompasses a range of engineering systems such as traffic networks and resource allocation. It describes the behavior of a group of agents who share a common set of $F$ facil…

A Case of Exponential Convergence Rates for SVM

2022-05-20 · Vivien Cabannes, Stefano Vigogna

Classification is often the first problem described in introductory machine learning classes. Generalization guarantees of classification have historically been offered by Vapnik-Chervonenkis theory. Yet those guarantees…

Classification