paper-with-me

홈 › Papers

SimQ-NAS: Simultaneous Quantization Policy and Neural Architecture Search

2023-12-19 · Sharath Nittur Sridhar, Maciej Szankin, Fang Chen, Sairam Sundaresan, Anthony Sarah

Recent one-shot Neural Architecture Search algorithms rely on training a hardware-agnostic super-network tailored to a specific task and then extracting efficient sub-networks for different hardware platforms. Popular approaches separate the training of super-networks from the search for sub-networks, often employing predictors to alleviate the computational overhead associated with search. Additionally, certain methods also incorporate the quantization policy within the search space. However, while the quantization policy search for convolutional neural networks is well studied, the extension of these methods to transformers and especially foundation models remains under-explored. In this paper, we demonstrate that by using multi-objective search algorithms paired with lightly trained predictors, we can efficiently search for both the sub-network architecture and the corresponding quantization policy and outperform their respective baselines across different performance objectives such as accuracy, model size, and latency. Specifically, we demonstrate that our approach performs well across both uni-modal (ViT and BERT) and multi-modal (BEiT-3) transformer-based architectures as well as convolutional architectures (ResNet). For certain networks, we demonstrate an improvement of up to $4.80x$ and $3.44x$ for latency and model size respectively, without degradation in accuracy compared to the fully quantized INT8 baselines.

📄 PDF Abstract BibTeX arXiv:2312.13301

Code (0)

등록된 구현이 없습니다.

Tasks

Neural Architecture SearchQuantization

Similar Papers 제목 키워드 기반

SimQFL: A Quantum Federated Learning Simulator with Real-Time Visualization

2025-08-17 · Ratun Rahman, Atit Pokharel, Md Raihan Uddin, Dinh C. Nguyen arxiv

Quantum federated learning (QFL) is an emerging field that has the potential to revolutionize computation by taking advantage of quantum physics concepts in a distributed machine learning (ML) environment. However, the m…

Federated Learning

LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference

2024-06-28 · Dong Liu, Yanxuan Yu

As large language models (LLMs) grow in size and deployment scale, quantization has become an essential technique for reducing memory footprint and improving inference efficiency. However, existing quantization toolkits …

GPUQuantization

QSTS: A Question-Sensitive Text Similarity Measure for Question Generation

2022-10-01 · COLING 2022 10 · Sujatha Das Gollapalli, See-Kiong Ng

While question generation (QG) has received significant focus in conversation modeling and text generation research, the problems of comparing questions and evaluation of QG models have remained inadequately addressed. I…

Question GenerationQuestion-GenerationQuestion SimilaritySemantic Similarity+3

HQNAS: Auto CNN deployment framework for joint quantization and architecture search

2022-10-16 · Hongjiang Chen, Yang Wang, Leibo Liu, Shaojun Wei 외

Deep learning applications are being transferred from the cloud to edge with the rapid development of embedded computing systems. In order to achieve higher energy efficiency with the limited resource budget, neural netw…

GPUNeural Architecture SearchQuantization

APQ: Joint Search for Network Architecture, Pruning and Quantization Policy

2020-06-15 · CVPR 2020 6 · Tianzhe Wang, Kuan Wang, Han Cai, Ji Lin 외

We present APQ for efficient deep learning inference on resource-constrained hardware. Unlike previous methods that separately search the neural architecture, pruning policy, and quantization policy, we optimize them in …

GPUQuantization