paper-with-me

홈 › Papers

EFloat: Entropy-coded Floating Point Format for Compressing Vector Embedding Models

2021-02-04 · NeurIPS 2021 12 · Rajesh Bordawekar, Bulent Abali, Ming-Hung Chen

In a large class of deep learning models, including vector embedding models such as word and database embeddings, we observe that floating point exponent values cluster around a few unique values, permitting entropy based data compression. Entropy coding compresses fixed-length values with variable-length codes, encoding most probable values with fewer bits. We propose the EFloat compressed floating point number format that uses a variable field boundary between the exponent and significand fields. EFloat uses entropy coding on exponent values and signs to minimize the average width of the exponent and sign fields, while preserving the original FP32 exponent range unchanged. Saved bits become part of the significand field increasing the EFloat numeric precision by 4.3 bits on average compared to other reduced-precision floating point formats. EFloat makes 8-bit and even smaller floats practical without sacrificing the exponent range of a 32-bit floating point representation. We currently use the EFloat format for saving memory capacity and bandwidth consumption of large vector embedding models such as those used for database embeddings. Using the RMS error as metric, we demonstrate that EFloat provides higher accuracy than other floating point formats with equal bit budget. The EF12 format with 12-bit budget has less end-to-end application error than the 16-bit BFloat16. EF16 with 16-bit budget has an RMS-error 17 to 35 times less than BF16 RMS-error for a diverse set of embedding models. When making similarity and dissimilarity queries, using the NDCG ranking metric, EFloat matches the result quality of prior floating point representations with larger bit budgets.

📄 PDF Abstract BibTeX arXiv:2102.02705

Code (0)

등록된 구현이 없습니다.

Tasks

Data Compression

Similar Papers 제목 키워드 기반

Exploring Approximations for Floating-Point Arithmetic using UppSAT

2017-11-24 · Aleksandar Zeljic, Peter Backeman, Christoph M. Wintersteiger, Philipp Ruemmer

We consider the problem of solving floating-point constraints obtained from software verification. We present UppSAT --- a new implementation of a systematic approximation refinement framework [ZWR17] as an abstract SMT …

To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration

2025-10-03 · Zeyu Yang, Tianyi Zhang, Jianwen Xie, Chuan Li 외 arxiv

The scaling of Generative AI (GenAI) models into the hundreds of billions of parameters makes low-precision computation indispensable for efficient deployment. We argue that the fundamental solution lies in developing lo…

Lossless Compression of Neural Network Components: Weights, Checkpoints, and K/V Caches in Low-Precision Formats

2025-08-20 · Anat Heilper, Doron Singer arxiv

As deep learning models grow and deployment becomes more widespread, reducing the storage and transmission costs of neural network weights has become increasingly important. While prior work such as ZipNN has shown that …

Improving Neural Network Efficiency via Post-Training Quantization With Adaptive Floating-Point

2021-01-01 · ICCV 2021 10 · Fangxin Liu, Wenbo Zhao, Zhezhi He, Yanzhi Wang 외

Model quantization has emerged as a mandatory technique for efficient inference with advanced Deep Neural Networks (DNN). It converts the model parameters in full precision (32-bit floating point) to the hardware fri…

Model CompressionQuantization

Theoretical Analysis of Relative Errors in Gradient Computations for Adversarial Attacks with CE Loss

2025-07-30 · Yunrui Yu, Hang Su, Cheng-zhong Xu, Zhizhong Su 외 arxiv

Gradient-based adversarial attacks using the Cross-Entropy (CE) loss often suffer from overestimation due to relative errors in gradient computation induced by floating-point arithmetic. This paper provides a rigorous th…