paper-with-me

홈 › Papers

ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention

2026-05-21 · Joe Sharratt arxiv

Efficient attention algorithms are critical to mitigate the quadratic cost of attention in long-context workloads. Prior work utilises block-scaled quantisation techniques on Blackwell GPUs to move attention computation to 4-bit precision to accelerate inference. However, these techniques result in significant quality degradation in long-context settings. We show that the output impact of quantisation error is highly non-uniform and increases with the importance of each query-key interaction, concentrating functionally relevant error in a small number of attention blocks that contain the most important tokens. We propose ThriftAttention, a low-bit attention variant that delivers near-FP16 long-context quality at FP4 inference efficiency. This approach proceeds in two stages. First, a heuristic rapidly selects a small number of important query-key block pairs for FP16 precision. Second, the selected blocks are computed in FP16 and the remaining blocks in FP4, with both paths merged via online softmax into a single output. We demonstrate across long-context benchmarks and model families that by computing only 5% of query-key blocks in FP16, ThriftAttention recovers on average 89.1% of the FP4-to-FP16 performance gap. We show ThriftAttention's advantage grows with sequence length, mitigating the systematic FP4 quality degradation observed at longer contexts. The code is available at https://github.com/joesharratt1229/ThriftAttention.

📄 PDF Abstract BibTeX arXiv:2605.23081

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization

2025-04-11 · Daeun Kim, Jinwoo Hwang, Changhun Oh, Jongse Park

Diffusion Transformer (DiT) has driven significant progress in image generation tasks. However, DiT inferencing is notoriously compute-intensive and incurs long latency even on datacenter-scale GPUs, primarily due to its…

Image GenerationQuantization

Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM Inference

2025-03-30 · Wei Tao, Bin Zhang, Xiaoyang Qu, Jiguang Wan 외

Recently, large language models (LLMs) have been able to handle longer and longer contexts. However, a context that is too long may cause intolerant inference latency and GPU memory usage. Existing methods propose mixed-…

GPUQuantization

MPX: Mixed Precision Training for JAX

2025-07-04 · Alexander Gräfe, Sebastian Trimpe arxiv

Mixed-precision training has emerged as an indispensable tool for enhancing the efficiency of neural network training in recent years. Concurrently, JAX has grown in popularity as a versatile machine learning toolbox. Ho…

Neural Quantum States in Mixed Precision

2026-01-28 · Massimo Solinas, Agnes Valenti, Nawaf Bou-Rabee, Roeland Wiersema arxiv

Scientific computing has long relied on double precision (64-bit floating point) arithmetic to guarantee accuracy in simulations of real-world phenomena. However, the growing availability of hardware accelerators such as…

Progressive Mixed-Precision Decoding for Efficient LLM Inference

2024-10-17 · Hao Mark Chen, Fuwen Tan, Alexandros Kouris, Royson Lee 외

In spite of the great potential of large language models (LLMs) across various tasks, their deployment on resource-constrained devices remains challenging due to their excessive computational and memory demands. Quantiza…

Quantization