paper-with-me

홈 › Papers

Power-of-Two Quantization-Aware-Training (PoT-QAT) in Large Language Models (LLMs)

2026-01-05 · Mahmoud Elgenedy arxiv

In Large Language Models (LLMs), the number of parameters has grown exponentially in the past few years, e.g., from 1.5 billion parameters in GPT-2 to 175 billion in GPT-3 to possibly more than trillion in higher versions. This raises a significant challenge for implementation, especially for Edge devices. Unlike cloud computing, memory and processing power for Edge devices are very limited, which necessitates developing novel ideas to make such applications feasible. In this work, we investigate compressing weights with a special quantization that limits numbers to only power-of-two (PoT). This helps save a huge amount of memory as only exponents need to be stored, more importantly, it significantly reduces processing power by replacing costly multiplication with low cost bit shifting. To overcome performance loss due to this strict quantization, we investigate Quantization Aware Training (QAT) to enhance performance through additional training. Results on GPT-2 124M show a major enhancement for quantized PoT model after additional training, with a perplexity enhancement of 66% and BERT-Score loss to baseline GPT-2 of 1%. The memory saving is estimated to be 87.5% while the inference speed is expected to be 3-10x faster with PoT quantization versus full-precision.

📄 PDF Abstract BibTeX arXiv:2601.02298

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques

2024-11-09 · Jahid Hasan

This paper presents a comprehensive analysis of quantization techniques for optimizing Large Language Models (LLMs), specifically focusing on Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT). Throug…

Quantization

Accumulator-Aware Post-Training Quantization

2024-09-25 · Ian Colbert, Fabian Grob, Giuseppe Franco, Jinjie Zhang 외

Several recent studies have investigated low-precision accumulation, reporting improvements in throughput, power, and area across various platforms. However, the accompanying proposals have only considered the quantizati…

image-classificationImage ClassificationQuantizationText Generation

LLMPi: Optimizing LLMs for High-Throughput on Raspberry Pi

2025-04-02 · Mahsa Ardakani, Jinendra Malekar, Ramtin Zand

Deploying Large Language Models (LLMs) on resource-constrained edge devices like the Raspberry Pi presents challenges in computational efficiency, power consumption, and response latency. This paper explores quantization…

Computational EfficiencyQuantization

How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training

2024-04-25 · Jaeseong You, Minseop Park, Kyunggeun Lee, Seokjun An 외

This paper investigates three different parameterizations of asymmetric uniform quantization for quantization-aware training: (1) scale and offset, (2) minimum and maximum, and (3) beta and gamma. We perform a comprehens…

Quantization

BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks

2024-06-24 · Jacob Nielsen, Peter Schneider-Kamp

Recently proposed methods for 1-bit and 1.58-bit quantization aware training investigate the performance and behavior of these methods in the context of large language models, finding state-of-the-art performance for mod…

Quantization