paper-with-me

홈 › Papers

Jumping through Local Minima: Quantization in the Loss Landscape of Vision Transformers

2023-08-21 · ICCV 2023 1 · Natalia Frumkin, Dibakar Gope, Diana Marculescu

Quantization scale and bit-width are the most important parameters when considering how to quantize a neural network. Prior work focuses on optimizing quantization scales in a global manner through gradient methods (gradient descent \& Hessian analysis). Yet, when applying perturbations to quantization scales, we observe a very jagged, highly non-smooth test loss landscape. In fact, small perturbations in quantization scale can greatly affect accuracy, yielding a $0.5-0.8\%$ accuracy boost in 4-bit quantized vision transformers (ViTs). In this regime, gradient methods break down, since they cannot reliably reach local minima. In our work, dubbed Evol-Q, we use evolutionary search to effectively traverse the non-smooth landscape. Additionally, we propose using an infoNCE loss, which not only helps combat overfitting on the small calibration dataset ($1,000$ images) but also makes traversing such a highly non-smooth surface easier. Evol-Q improves the top-1 accuracy of a fully quantized ViT-Base by $10.30\%$, $0.78\%$, and $0.15\%$ for $3$-bit, $4$-bit, and $8$-bit weight quantization levels. Extensive experiments on a variety of CNN and ViT architectures further demonstrate its robustness in extreme quantization scenarios. Our code is available at https://github.com/enyac-group/evol-q

📄 PDF Abstract BibTeX arXiv:2308.10814

Code (1)

enyac-group/evol-q 공식 구현 pytorch

Tasks

Quantization

Methods 이 논문이 사용한 방법론

InfoNCE 설명 없음

Similar Papers 제목 키워드 기반

Implicit bias of SGD in $L_{2}$-regularized linear DNNs: One-way jumps from high to low rank

2023-05-25 · Zihan Wang, Arthur Jacot

The $L_{2}$-regularized loss of Deep Linear Networks (DLNs) with more than one hidden layers has multiple local minima, corresponding to matrices with different ranks. In tasks such as matrix completion, the goal is to c…

Matrix Completion

SQuAT: Sharpness- and Quantization-Aware Training for BERT

2022-10-13 · Zheng Wang, Juncheng B Li, Shuhui Qu, Florian Metze 외

Quantization is an effective technique to reduce memory footprint, inference latency, and power consumption of deep learning models. However, existing quantization methods suffer from accuracy degradation compared to ful…

Quantization

VecQ: Minimal Loss DNN Model Compression With Vectorized Weight Quantization

2020-05-18 · Cheng Gong, Yao Chen, Ye Lu, Tao Li 외

Quantization has been proven to be an effective method for reducing the computing and/or storage cost of DNNs. However, the trade-off between the quantization bitwidth and final accuracy is complex and non-convex, which …

Model Compressionobject-detectionObject DetectionQuantization

Autonomous Navigation for Quadrupedal Robots with Optimized Jumping through Constrained Obstacles

2021-07-01 · Scott Gilroy, Derek Lau, Lizhi Yang, Ed Izaguirre 외

Quadrupeds are strong candidates for navigating challenging environments because of their agile and dynamic designs. This paper presents a methodology that extends the range of exploration for quadrupedal robots by creat…

Autonomous NavigationDecision MakingNavigate

QT-DoG: Quantization-aware Training for Domain Generalization

2024-10-08 · Saqib Javed, Hieu Le, Mathieu Salzmann

Domain Generalization (DG) aims to train models that perform well not only on the training (source) domains but also on novel, unseen target data distributions. A key challenge in DG is preventing overfitting to source d…

Domain GeneralizationModel CompressionQuantization