paper-with-me

Papers

HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization

2026-05-28 · Artur Zagitov, Gleb Molodtsov, Aleksandr Beznosikov arxiv

Post-training quantization (PTQ) is essential for deploying LLMs under memory and bandwidth constraints. However, extreme low-bit quantization remains highly sensitive to activation outliers and anisotropic weight curvature. Existing incoherence-based PTQ methods mitigate this issue with fixed randomized Hadamard transforms (RHTs), which improve quantization robustness but cannot adapt the rotated basis to the layer, calibration distribution, or quantizer. We introduce HARP (Hadamard-preconditioned Adaptive Rotation Processor), a learnable structured two-sided orthogonal processor that replaces fixed Hadamard mixing while preserving exact full-precision equivalence. HARP represents each rotation as a product of sparse butterfly-like block-orthogonal stages, supports non-power-of-two dimensions via Mixed-Radix schedules, and initializes to the RHT processor up to a fixed permutation. Fitted only on calibration data, HARP adapts the quantization basis to each layer and backend. Across 2-4 bit settings on models ranging from 1B to 70B parameters, HARP improves perplexity and zero-shot accuracy over fixed RHT. Importantly, HARP preserves deployment efficiency, reaching 128 tok/s versus 61 tok/s for FP16.

📄 PDF Abstract BibTeX arXiv:2605.29843

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Approximating Uniform Random Rotations by Two-Block Structured Hadamard Rotations in High Dimensions

2026-04-25 · Tomer Zilca, Gal Mendelson arxiv

Uniform random rotations are a useful primitive in applications such as fast Johnson-Lindenstrauss embeddings, kernel approximation, communication-efficient learning, and recent AI compression pipelines, but they are com…

ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms

2025-09-11 · Bingxin Xu, Zhen Dong, Oussama Elachqar, Yuzhang Shang arxiv

Large language models require massive memory footprints, severely limiting deployment on consumer hardware. Quantization reduces memory through lower numerical precision, but extreme 2-bit quantization suffers from catas…

AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation

2026-04-02 · Seonggon Kim, Alireza Khodamoradi, Pranathi Vasireddy, Kristof Denolf 외 arxiv

Hadamard transforms have become a key tool for stabilizing low-precision training, but existing methods apply them uniformly across tensors and computation paths. We show that this one-size-fits-all strategy is inherentl…

PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression

2026-03-30 · Caio Vicentino arxiv

We present PolarQuant, a post-training weight quantization method for large language models (LLMs) that exploits the distributional structure of neural network weights to achieve near-lossless compression. PolarQuant ope…

WUSH: Near-Optimal Adaptive Transforms for LLM Quantization

2025-11-30 · Jiale Chen, Vage Egiazarian, Roberto L. Castro, Torsten Hoefler 외 arxiv

Quantizing LLM weights and activations is a standard approach for efficient deployment, but a few extreme outliers can stretch the dynamic range and amplify low-bit quantization errors. Prior transform-based mitigations …