paper-with-me

Papers

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

2026-08-21 · Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús arxiv

Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is never independently trained at full precision, its bfloat16 checkpoint is a distillation-recovered approximation of the original; QAH distills the 4-bit student directly from the original, uncompressed model. On a GPT-OSS 120B to 60B to MXFP4 pipeline, the QAH student matches or beats its bfloat16 source on 7 of 9 benchmarks at roughly 4 times less weight memory and half the teacher's parameter count, and is released open-weight as Hypernova-60B. Against a matched QAT baseline it reaches a comparable peak about 7 times faster and stays stable under continued training, without hand-tuned early stopping. We also report deployment lessons, including a large, reproducible quality gap between distributed-training backends. Our aim is a recipe deployable without a multi-week hyper-parameter search.

📄 PDF Abstract BibTeX arXiv:2608.20953

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Not All NVFP4 QAT Recipes Are Equal: How Architecture and Scale Shape Model Quality for Anomaly Segmentation

2026-05-26 · Zijian Du, Oleg Rybakov arxiv

Real-time anomaly segmentation demands both high recall and efficient low-precision inference. We study the three-way interaction of model architecture, model scale, and FP4 quantization-aware training (QAT) recipe on a …

Brain Tumor Segmentation

QuantX: A Framework for Hardware-Aware Quantization of Generative AI Workloads

2025-05-12 · Khurram Mazher, Saad Bin Nasir

We present QuantX: a tailored suite of recipes for LLM and VLM quantization. It is capable of quantizing down to 3-bit resolutions with minimal loss in performance. The quantization strategies in QuantX take into account…

Quantization

Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models

2026-08-21 · Zhen Yang, Sizai Hou, Kaiwen Zheng, Yaofang Liu 외 arxiv

Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data select…

FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error

2025-11-04 · Fengjuan Wang, Zhiyi Su, Xingzhu Hu, Cheng Wang 외 arxiv

Training large Mixture-of-Experts (MoE) models remains computationally prohibitive due to their extreme compute and memory demands. Although low-precision training promises to accelerate computation and reduce memory foo…

A Speed Odyssey for Deployable Quantization of LLMs

2023-11-16 · Qingyuan Li, Ran Meng, Yiduo Li, Bo Zhang 외

The large language model era urges faster and less costly inference. Prior model compression works on LLMs tend to undertake a software-centric approach primarily focused on the simulated quantization performance. By neg…

Language ModelingLanguage ModellingLarge Language ModelModel Compression+1