Extreme Model Compression for On-device Natural Language Understanding
In this paper, we propose and experiment with techniques for extreme compression of neural natural language understanding (NLU) models, making them suitable for execution on resource-constrained devices. We propose a task-aware, end-to-end compression approach that performs word-embedding compression jointly with NLU task learning. We show our results on a large-scale, commercial NLU system trained on a varied set of intents with huge vocabulary sizes. Our approach outperforms a range of baselines and achieves a compression rate of 97.4% with less than 3.7% degradation in predictive performance. Our analysis indicates that the signal from the downstream task is important for effective compression with minimal degradation in performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Model CompressionNatural Language UnderstandingSimilar Papers 제목 키워드 기반
Aggressive Post-Training Compression on Extremely Large Language Models
The increasing size and complexity of Large Language Models (LLMs) pose challenges for their deployment on personal computers and mobile devices. Aggressive post-training model compression is necessary to reduce the mode…
Model CompressionNetwork PruningQuantizationStatistical Model Compression for Small-Footprint Natural Language Understanding
In this paper we investigate statistical model compression applied to natural language understanding (NLU) models. Small-footprint NLU models are important for enabling offline systems on hardware restricted devices, and…
Model CompressionNatural Language UnderstandingQuantizationDevice Tuning for Multi-Task Large Model
Unsupervised pre-training approaches have achieved great success in many fields such as Computer Vision (CV), Natural Language Processing (NLP) and so on. However, compared to typical deep learning models, pre-training o…
modelMulti-Task LearningUnsupervised Pre-trainingExtreme Compression of Large Language Models via Additive Quantization
The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices. In this paper, we revisit the problem of "ext…
CPUGPUInformation RetrievalQuantizationMap-Assisted Remote-Sensing Image Compression at Extremely Low Bitrates
Remote-sensing (RS) image compression at extremely low bitrates has always been a challenging task in practical scenarios like edge device storage and narrow bandwidth transmission. Generative models including VAEs and G…
Image Compression