Papers Data Free Quantization
“Data Free Quantization” 태그가 달린 논문 37편 · 필터 해제
DVD-Quant: Data-free Video Diffusion Transformers Quantization
Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture for video generation, yet their computational and memory demands hinder practical deployment. While post-training quantization (PTQ) present…
Data Free QuantizationQuantizationVideo GenerationOuroMamba: A Data-Free Quantization Framework for Vision Mamba Models
We present OuroMamba, the first data-free post-training quantization (DFQ) method for vision Mamba-based models (VMMs). We identify two key challenges in enabling DFQ for VMMs, (1) VMM's recurrent state transitions restr…
channel selectionContrastive LearningData Free QuantizationGPU+3Enhancing Diversity for Data-free Quantization
Model quantization is an effective way to compress deep neural networks and accelerate the inference time on edge devices. Existing quantization methods usually require original data for calibration during the compre…
Data Free QuantizationDiversityQuantizationSemantics Prompting Data-Free Quantization for Low-Bit Vision Transformers
Data-free quantization (DFQ), which facilitates model quantization without real data to address increasing concerns about data security, has garnered significant attention within the model compression community. Recently…
Data Free QuantizationModel CompressionQuantizationRelation-Guided Adversarial Learning for Data-free Knowledge Transfer
Data-free knowledge distillation transfers knowledge by recovering training data from a pre-trained model. Despite the recent success of seeking global data diversity, the diversity within each class and the similarity a…
Data-free Knowledge DistillationData Free QuantizationDiversityImage Generation+6A method of using RSVD in residual calculation of LowBit GEMM
The advancements of hardware technology in recent years has brought many possibilities for low-precision applications. However, the use of low precision can introduce significant computational errors, posing a considerab…
Data Free QuantizationQuantizationPrivacy-Preserving SAM Quantization for Efficient Edge Intelligence in Healthcare
The disparity in healthcare personnel expertise and medical resources across different regions of the world is a pressing social issue. Artificial intelligence technology offers new opportunities to alleviate this issue.…
Data Free QuantizationImage SegmentationModel CompressionPrivacy Preserving+2MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
Data-free quantization (DFQ) is a technique that creates a lightweight network from its full-precision counterpart without the original training data, often through a synthetic dataset. Although several DFQ methods have …
Data Free QuantizationQuantizationEasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
Large language models (LLMs) have proven to be very superior to conventional methods in various tasks. However, their expensive computations and high memory requirements are prohibitive for deployment. Model quantization…
Data Free QuantizationQuantizationData-Free Quantization via Pseudo-label Filtering
Quantization for model compression can efficiently reduce the network complexity and storage requirement but the original training data is necessary to remedy the performance loss caused by quantization. The Data-Fre…
Data Free QuantizationModel CompressionPseudo LabelPseudo Label Filtering+1Robustness-Guided Image Synthesis for Data-Free Quantization
Quantization has emerged as a promising direction for model compression. Recently, data-free quantization has been widely studied as a promising method to avoid privacy concerns, which synthesizes images as an alternativ…
Data Free QuantizationDiversityImage GenerationModel Compression+1Causal-DFQ: Causality Guided Data-free Network Quantization
Model quantization, which aims to compress deep neural networks and accelerate inference speed, has greatly facilitated the development of cumbersome models on mobile and edge devices. There is a common assumption in qua…
Data Free QuantizationNeural Network CompressionQuantizationData-Free Quantization via Mixed-Precision Compensation without Fine-Tuning
Neural network quantization is a very promising solution in the field of model compression, but its resulting accuracy highly depends on a training/fine-tuning process and requires the original data. This not only brings…
Data Free QuantizationModel CompressionQuantizationQNNRepair: Quantized Neural Network Repair
We present QNNRepair, the first method in the literature for repairing quantized neural networks (QNNs). QNNRepair aims to improve the accuracy of a neural network model after quantization. It accepts the full-precision …
Data Free QuantizationFault localizationQuantizationTowards Accurate Post-training Quantization for Diffusion Models
In this paper, we propose an accurate data-free post-training quantization framework of diffusion models (ADP-DM) for efficient image generation. Conventional data-free quantization methods learn shared quantization func…
Data Free QuantizationImage GenerationQuantizationLLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Several post-training quantization methods have been applied to large language models (LLMs), and have been shown to perform well down to 8-bits. We find that these methods break down at lower bit precision, and investig…
Data Free QuantizationQuantizationAdaptive Data-Free Quantization
Data-free quantization (DFQ) recovers the performance of quantized network (Q) without the original data, but generates the fake sample via a generator (G) by learning from full-precision network (P), which, however, is …
Data Free QuantizationQuantizationRethinking Data-Free Quantization as a Zero-Sum Game
Data-free quantization (DFQ) recovers the performance of quantized network (Q) without accessing the real data, but generates the fake sample via a generator (G) by learning from full-precision network (P) instead. Howev…
Data Free QuantizationQuantizationACQ: Improving Generative Data-free Quantization Via Attention Correction
Data-free quantization aims to achieve model quantization without accessing any authentic sample. It is significant in an application-oriented context involving data privacy. Converting noise vectors into synthetic sampl…
Data Free QuantizationPositionQuantizationGenie: Show Me the Data for Quantization
Zero-shot quantization is a promising approach for developing lightweight deep neural networks when data is inaccessible owing to various reasons, including cost and issues related to privacy. By exploiting the learned p…
Data Free QuantizationQuantization