paper-with-me

홈 › Papers

OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension

2026-04-14 · Zhiyuan Zhang, Yanzhao Li, Zhiqiang Zou, Bai Du, Yupeng Sun, Hui Dong, Hui Wang arxiv

While 4-bit quantization is essential for high-throughput deployment of Large Language Models, activation outliers often lead to significant accuracy degradation due to the restricted dynamic range of low-bit formats. In this paper, we systematically investigate the spatial distribution of outliers and demonstrate a token-persistent structural clustering effect, where high-magnitude outliers consistently occupy fixed channels across tokens. Building on this insight, we propose OSC, a hardware-efficient framework for outlier suppression. During inference, OSC executes a dual-path computation consisting of a low-precision 4-bit General Matrix Multiplication (GEMM) path and a high-precision 16-bit branch GEMM path. Specifically, OSC uses an offline group-wise strategy to identify the channels where outliers are located and then performs structured sub-tensor extraction to coalesce these scattered activation channels into a compact dense tensor online. This mechanism implements outlier protection through regularized and high-throughput GEMM operations, achieving a seamless fit with modern 4-bit micro-scaling hardware. Furthermore, for the inputs of W2 where outlier clustering is less pronounced, we integrate a fallback strategy to FP8. Evaluation on Qwen3-8B and Qwen3-30B restricts the average accuracy drop to 2.19 and 1.12 points, respectively. Notably, OSC is highly hardware-friendly, achieving a peak speedup of 1.78x over the W8A8 GEMM baseline on a modern AI accelerator.

📄 PDF Abstract BibTeX arXiv:2604.12782

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Improving Neural Network Quantization without Retraining using Outlier Channel Splitting

2019-01-28 · Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Christopher De Sa 외

Quantization can improve the execution latency and energy efficiency of neural networks on both commodity GPUs and specialized accelerators. The majority of existing literature focuses on training quantized DNNs, while t…

Language ModelingLanguage ModellingNeural Network CompressionQuantization

OutlierTune: Efficient Channel-Wise Quantization for Large Language Models

2024-06-27 · Jinguang Wang, Yuexi Yin, Haifeng Sun, Qi Qi 외

Quantizing the activations of large language models (LLMs) has been a significant challenge due to the presence of structured outliers. Most existing methods focus on the per-token or per-tensor quantization of activatio…

Quantization

Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models

2023-09-27 · Jung Hwan Heo, Jeonghoon Kim, Beomseok Kwon, Byeongwook Kim 외

Large Language Models (LLMs) have recently demonstrated remarkable success across various tasks. However, efficiently serving LLMs has been a challenge due to the large memory bottleneck, specifically in small batch infe…

HumanEvalLanguage ModelingLanguage ModellingMMLU+1

MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition

2026-04-06 · Seoungsub Lee, In Seo Kim, Seon Wook Kim arxiv

Large language models (LLMs) have achieved outstanding performance across a wide range of natural language processing tasks, but their enormous parameter counts impose ubstantial memory and computational overheads. This …

ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA

2026-05-03 · Shengzhe Lyu, Yuhan She, Patrick S. Y. Hung, Ray C. C. Cheung 외 arxiv

Vision Mamba (ViM) models offer a compelling efficiency advantage over Transformers by leveraging the linear complexity of State Space Models (SSMs), yet efficiently deploying them on FPGAs remains challenging. Linear la…