paper-with-me

홈 › Papers

SageAttention2++: A More Efficient Implementation of SageAttention2

2025-05-27 · Jintao Zhang, Xiaoming Xu, Jia Wei, Haofeng Huang, Pengle Zhang, Chendong Xiang, Jun Zhu, Jianfei Chen

The efficiency of attention is critical because its time complexity grows quadratically with sequence length. SageAttention2 addresses this by utilizing quantization to accelerate matrix multiplications (Matmul) in attention. To further accelerate SageAttention2, we propose to utilize the faster instruction of FP8 Matmul accumulated in FP16. The instruction is 2x faster than the FP8 Matmul used in SageAttention2. Our experiments show that SageAttention2++ achieves a 3.9x speedup over FlashAttention while maintaining the same attention accuracy as SageAttention2. This means SageAttention2++ effectively accelerates various models, including those for language, image, and video generation, with negligible end-to-end metrics loss. The code will be available at https://github.com/thu-ml/SageAttention.

📄 PDF Abstract BibTeX arXiv:2505.21136

Code (2)

thu-ml/SageAttention 공식 구현 pytorch
thu-ml/spargeattn pytorch

Tasks

QuantizationVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization

2024-11-17 · Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei 외

Although quantization for linear layers has been widely used, its application to accelerate the attention process remains limited. To further enhance the efficiency of attention computation compared to SageAttention whil…

Image GenerationQuantizationVideo Generation

SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training

2025-05-16 · Jintao Zhang, Jia Wei, Pengle Zhang, Xiaoming Xu 외

The efficiency of attention is important due to its quadratic time complexity. We enhance the efficiency of attention through two key contributions: First, we leverage the new FP4 Tensor Cores in Blackwell GPUs to accele…

SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration

2024-10-03 · Jintao Zhang, Jia Wei, Haofeng Huang, Pengle Zhang 외

The transformer architecture predominates across various models. As the heart of the transformer, attention has a computational complexity of O(N^2), compared to O(N) for linear transformations. When handling large seque…

Image GenerationQuantizationVideo Generation

SageBwd: A Trainable Low-bit Attention

2026-03-02 · Jintao Zhang, Marco Chen, Haoxu Wang, Kai Jiang 외 arxiv

Low-bit attention, such as SageAttention, has emerged as an effective approach for accelerating model inference, but its applicability to training remains poorly understood. In prior work, we introduced SageBwd, a traina…

TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times

2025-12-18 · Jintao Zhang, Kaiwen Zheng, Kai Jiang, Haoxu Wang 외 arxiv

We introduce TurboDiffusion, a video generation acceleration framework that can speed up end-to-end diffusion generation by 100-200x while maintaining video quality. TurboDiffusion mainly relies on several components for…

Video Generation