paper-with-me

Papers

H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference

2026-08-28 · Hao Yu, Zheng Li, Dayiheng Liu, Jianwei Zhang arxiv

The NVIDIA Blackwell architecture, with native support for the ultra-fine-grained NVFP4 format, opens new opportunities for accelerating large language model (LLM) inference. NVFP4's micro-block design, such as a group size of 16, offers strong representational flexibility for capturing local weight distributions and isolating outliers, but it also introduces a large and highly sensitive space of per-group scaling factors. Existing post-training quantization (PTQ) methods primarily focus on refining quantized weight values, leaving this scale-selection step underexplored. To address this gap, we propose \textbf{H-Scale}, a lightweight post-processing method for NVFP4 per-group scale refinement. Instead of minimizing plain weight reconstruction error, H-Scale selects hardware-valid group scales using a diagonal second-order proxy derived from calibration activations, thereby targeting layer output perturbation more directly. It is designed as a drop-in replacement for RTN-style scale selection in diverse NVFP4 pipelines, requires only modest offline calibration, and introduces strictly zero overhead at inference time. Under a fixed evaluation protocol, experiments on mainstream LLMs show that H-Scale generally improves a broad range of NVFP4 baselines and brings several variants closer to the BF16 reference.

📄 PDF Abstract BibTeX arXiv:2608.28113

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization

2026-05-12 · Chengzhu Bao, Xianglong Yan, Zhiteng Li, Guangshuo Qin 외 arxiv

NVFP4 has recently emerged as an efficient 4-bit microscaling format for large language models (LLMs), offering superior numerical fidelity with native hardware support. However, existing methods often yield suboptimal p…

ScaleSweep: Accurate NVFP4 Post-Training Quantization of LLMs via Block Scale Initialization

2026-05-30 · Li Lin, Xiaojun Wan arxiv

NVFP4 is a recently introduced hardware-supported FP4 format that improves the fidelity of 4-bit quantization through fine-grained block scales. However, existing NVFP4 scale initialization methods still primarily rely o…

Characterizing the Impact of NVFP4 Quantization for Low-Power Edge AI Deployment

2026-06-03 · Ovishake Sen, Venkata Nithin Kamineni, Daniel Lobo, Swarup Bhunia 외 arxiv

Energy-efficient neural-network inference at the edge requires reducing arithmetic cost, memory traffic, computation energy, and storage overhead while maintaining acceptable accuracy. This paper presents an ablation-foc…

Adaptive Block-Scaled Data Types

2026-03-30 · Jack Cook, Hyemin S. Lee, Kathryn Le, Junxian Guo 외 arxiv

NVFP4 has grown increasingly popular as a 4-bit format for quantizing large language models due to its hardware support and its ability to retain useful information with relatively few bits per parameter. However, the fo…

ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling

2026-06-11 · Sihwa Lee, Janghwan Lee, Donghoon Yoo, Jae Gon Kim 외 arxiv

Large reasoning models (LRMs) improve complex problem-solving by generating long intermediate reasoning traces, but this substantially increases inference costs. NVFP4 inference offers a promising approach to reduce both…