paper-with-me

Papers

IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference

2025-11-26 · Wanli Zhong, Haibo Feng, Zirui Zhou, Hanyang Peng, Shiqi Yu arxiv

Deploying Transformer models on edge devices is limited by latency and energy budgets. While INT8 quantization effectively accelerates the primary matrix multiplications, it exposes the softmax-related path as the dominant bottleneck. This stage incurs a costly dequantize -> softmax -> requantize detour, which can account for up to 65% of total attention latency and disrupts the end-to-end integer dataflow critical for edge hardware efficiency. To address this limitation, we present IntAttention, the first fully integer attention pipeline that serves as a training-free drop-in replacement. At the core of our approach lies IndexSoftmax, a hardware-friendly operator that replaces floating-point exponentials entirely within the integer domain. IntAttention integrates sparsity-aware clipping, a 32-entry lookup table approximation, and direct integer normalization, thereby eliminating datatype conversion overhead along the attention path. Experiments on Armv8 CPUs show that our method achieves up to 3.7x speedup and 61% energy reduction over FP16 baselines, and up to 2.0x speedup over conventional INT8 attention pipelines. Across diverse language and vision models, as well as additional reasoning and long-context evaluations, IntAttention maintains strong overall fidelity and demonstrates a more favorable trade-off than existing LUT-based softmax approximations. Code is available at https://github.com/WanliZhong/IntAttention

📄 PDF Abstract BibTeX arXiv:2511.21513

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Is Integer Arithmetic Enough for Deep Learning Training?

2022-07-18 · Alireza Ghaffari, Marzieh S. Tahaei, Mohammadreza Tayaranian, Masoud Asgharian 외

The ever-increasing computational complexity of deep learning models makes their training and deployment difficult on various cloud and edge platforms. Replacing floating-point arithmetic with low-bit integer arithmetic …

Deep Learningobject-detectionObject DetectionQuantization+1

Batch Normalization-Free Fully Integer Quantized Neural Networks via Progressive Tandem Learning

2025-12-18 · Pengfei Sun, Wenyu Jiang, Piew Yoong Chee, Paul Devos 외 arxiv

Quantised neural networks (QNNs) shrink models and reduce inference energy through low-bit arithmetic, yet most still depend on a running statistics batch normalisation (BN) layer, preventing true integer-only deployment…

Optimizing Explicit Unit-Distance Lower-Bound Certificates

2026-06-02 · Michael T. M. Emmerich arxiv

The 2026 disproof of Erdős's unit-distance conjecture and Sawin's quantitative refinement show that the maximum number $u(n)$ of unit distances among $n$ planar points can exceed $n^{1+\varepsilon}$ for a fixed positive …

Markov Logic Networks for Text Mining: A Qualitative and Empirical Comparison with Integer Linear Programming

2016-05-01 · LREC 2016 5 · Luis Gerardo Mojica de la Vega, Vincent Ng

Joint inference approaches such as Integer Linear Programming (ILP) and Markov Logic Networks (MLNs) have recently been successfully applied to many natural language processing (NLP) tasks, often outperforming their pipe…

I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

2024-05-28 · Xing Hu, Yuan Cheng, Dawei Yang, Zhihang Yuan 외

Post-training quantization (PTQ) serves as a potent technique to accelerate the inference of large language models (LLMs). Nonetheless, existing works still necessitate a considerable number of floating-point (FP) operat…

Quantization