Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks
In this work, we developed and tested 3 techniques for vector quantization (VQ) based model weight compression. To mitigate codebook collapse and enable end-to-end training, we adopted cosine similarity-based assignment. Building on ideas from attention-based formulations in Differentiable K-Means (DKM), we further improved this approach by using cosine similarity for assignment combined with top-1 sampling and a straight-through estimator, thereby eliminating the need for weighted-average reconstruction. Finally, we investigated the use of differentiable neural architecture search (NAS) to adaptively select layer-wise quantization configurations, further optimizing the compression process. Although our method does not consistently outperform existing approaches across all quantization levels, it provides useful insights into the design trade-offs and behaviors of VQ-based model compression methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Neural Architecture SearchModel CompressionSimilar Papers 제목 키워드 기반
Post-training Quantization with Multiple Points: Mixed Precision without Mixed Precision
We consider the post-training quantization problem, which discretizes the weights of pre-trained deep neural networks without re-training the model. We propose multipoint quantization, a quantization method that approxim…
object-detectionObject DetectionQuantizationCollaborative Automotive Radar Sensing via Mixed-Precision Distributed Array Completion
This paper investigates the effects of coarse quantization with mixed precision on measurements obtained from sparse linear arrays, synthesized by a collaborative automotive radar sensing strategy. The mixed quantization…
Matrix CompletionQuantizationOn the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note
This short note presents a dimension-independent subgaussian concentration bound for Gaussian vectors under coordinate-wise nonlinear mappings. Discovered by Gemini 3.5 Flash, this result applies to any bounded function …
Resource Allocation and Dithering of Bayesian Parameter Estimation Using Mixed-Resolution Data
Quantization of signals is an integral part of modern signal processing applications, such as sensing, communication, and inference. While signal quantization provides many physical advantages, it usually degrades the su…
parameter estimationQuantizationEffective and Efficient Mixed Precision Quantization of Speech Foundation Models
This paper presents a novel mixed-precision quantization approach for speech foundation models that tightly integrates mixed-precision learning and quantized model parameter estimation into one single model compression s…
Model Compressionparameter estimationQuantization