paper-with-me

홈 › Papers

Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference

2026-04-05 · Yifu Ding, Xinhao Zhang, Jinyang Guo arxiv

Transformer-based large language models (LLMs) have demonstrated remarkable performance across a wide range of real-world tasks, but their inference cost remains prohibitively high due to the quadratic complexity of attention and the memory bandwidth limitations of high-precision operations. In this work, we present a low-bit mixed-precision attention kernel using the microscaling floating-point (MXFP) data format, utilizing the computing capability on next-generation GPU architectures. Our Diagonal-Tiled Mixed-Precision Attention (DMA) incorporates two kinds of low-bit computation at the tiling-level, and is a delicate fused kernel implemented using Triton, exploiting hardware-level parallelism and memory efficiency to enable fast and efficient inference without compromising model performance. Extensive empirical evaluations on NVIDIA B200 GPUs show that our kernel maintains generation quality with negligible degradation, and meanwhile achieves significant speedup by kernel fusion. We release our code at https://github.com/yifu-ding/MP-Sparse-Attn.

📄 PDF Abstract BibTeX arXiv:2604.03950

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models

2025-08-04 · Wenyuan Liu, Haoqian Meng, Yilun Luo, Yafei Zhao 외 arxiv

Quantization significantly accelerates inference in large language models (LLMs) by replacing original high-precision matrices with low-precision counterparts. Recent advances in weight-activation quantization have prima…

Mathematical ReasoningCode Generation

SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs

2025-12-04 · Wenhua Cheng, Weiwei Zhang, Heng Guo, Haihao Shen 외 arxiv

Extremely low-bit quantization is critical for efficiently deploying Large Language Models (LLMs), yet it often leads to severe performance degradation at 2 bits and even at 4 bits (e.g., MXFP4). We present SignRoundV2, …

TiledAttention: a CUDA Tile SDPA Kernel for PyTorch

2026-03-02 · Taimur Khan arxiv

TiledAttention is a scaled dot-product attention (SDPA) forward operator for SDPA research on NVIDIA GPUs. Implemented in cuTile Python (TileIR) and exposed as a PyTorch-callable function, it is easier to modify than low…

W4A4 Quantization for Inference on Wan2.2-I2V-A14B

2026-06-28 · Yidong Chen, Chengyu Shi, Jiahao Liu arxiv

We summarize our submission to Sub-Challenge 1: W4A4 Quantization for Inference (HiF4 / MXFP4) of the ICME 2026 Low-Bit-width Large-Model Quantization Challenge. The sub-challenge targets 4-bit weight and 4-bit activatio…

Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration

2025-11-09 · Stef Cuyckens, Xiaoling Yi, Robin Geens, Joren Dumoulin 외 arxiv

Emerging continual learning applications necessitate next-generation neural processing unit (NPU) platforms to support both training and inference operations. The promising Microscaling (MX) standard enables narrow bit-w…

Continual Learning