paper-with-me

홈 › Papers

Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2

2025-12-29 · Yilun Luo, Huaqing Zheng, Haoqian Meng, Wenyuan Liu, Peng Zhang arxiv

Huawei's openPangu-Embedded-1B and openPangu-Embedded-7B are variants of the openPangu large language model, designed for efficient deployment on Ascend NPUs. The 7B variant supports three distinct Chain-of-Thought (CoT) reasoning paradigms, namely slow_think, auto_think, and no_think, while the 1B variant operates exclusively in the no_think mode, which employs condensed reasoning for higher efficiency. Although CoT reasoning enhances capability, the generation of extended reasoning traces introduces substantial memory and latency overheads, posing challenges for practical deployment on Ascend NPUs. This paper addresses these computational constraints by leveraging low-bit quantization, which transforms FP16 computations into more efficient integer arithmetic. We introduce a unified low-bit inference framework, supporting INT8 (W8A8) and W4A8 quantization, specifically optimized for openPangu-Embedded models on the Atlas A2. Our comprehensive evaluation on code generation benchmarks (HumanEval and MBPP) demonstrates the efficacy of this approach. INT8 quantization consistently preserves over 90\% of the FP16 baseline accuracy and achieves a 1.5x prefill speedup on the Atlas A2. Furthermore, W4A8 quantization significantly reduces memory consumption, albeit with a moderate trade-off in accuracy. These findings collectively indicate that low-bit quantization effectively facilitates efficient CoT reasoning on Ascend NPUs, maintaining high model fidelity.

📄 PDF Abstract BibTeX arXiv:2512.23367

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

An Empirical Study of OpenPangu Quantization on Ascend NPUs

2026-06-19 · Tong Shi, Jiacheng Wang, Hui Xie, Ying Li 외 arxiv

OpenPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs has not been systematically characterized.…

Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training

2026-05-25 · Yingying Cheng, Jinquan Shi, Li Zhou, Zhiyang He 외 arxiv

Quantization-aware training (QAT) with low-bit floating-point formats enables efficient LLM deployment, yet introduces subtle failure modes invisible to standard training metrics. We present a systematic study of HiF8 W8…

Loss Aware Post-training Quantization

2019-11-17 · Yury Nahshan, Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii 외

Neural network quantization enables the deployment of large models on resource-constrained devices. Current post-training quantization methods fall short in terms of accuracy for INT4 (or lower) but provide reasonable ac…

Quantization

E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring

2026-05-16 · Wenjun Wang, Yanggan Gu, Shuo Cai, Yuanyi Wang 외 arxiv

Low-resource deployment constraints have made model quantization essential for deploying neural networks while preserving performance. Meanwhile, model merging has become an increasingly practical low-resource strategy f…

Post-Training Piecewise Linear Quantization for Deep Neural Networks

2020-01-31 · ECCV 2020 8 · Jun Fang, Ali Shafiee, Hamzah Abdel-Aziz, David Thorsley 외

Quantization plays an important role in the energy-efficient deployment of deep neural networks on resource-limited devices. Post-training quantization is highly desirable since it does not require retraining or access t…

image-classificationImage Classificationobject-detectionObject Detection+2