paper-with-me

홈 › Papers

Extreme Model Compression for On-device Natural Language Understanding

2020-11-30 · COLING 2020 8 · Kanthashree Mysore Sathyendra, Samridhi Choudhary, Leah Nicolich-Henkin

In this paper, we propose and experiment with techniques for extreme compression of neural natural language understanding (NLU) models, making them suitable for execution on resource-constrained devices. We propose a task-aware, end-to-end compression approach that performs word-embedding compression jointly with NLU task learning. We show our results on a large-scale, commercial NLU system trained on a varied set of intents with huge vocabulary sizes. Our approach outperforms a range of baselines and achieves a compression rate of 97.4% with less than 3.7% degradation in predictive performance. Our analysis indicates that the signal from the downstream task is important for effective compression with minimal degradation in performance.

📄 PDF Abstract BibTeX arXiv:2012.00124

Code (0)

등록된 구현이 없습니다.

Tasks

Model CompressionNatural Language Understanding

Similar Papers 제목 키워드 기반

Aggressive Post-Training Compression on Extremely Large Language Models

2024-09-30 · Zining Zhang, Yao Chen, Bingsheng He, Zhenjie Zhang

The increasing size and complexity of Large Language Models (LLMs) pose challenges for their deployment on personal computers and mobile devices. Aggressive post-training model compression is necessary to reduce the mode…

Model CompressionNetwork PruningQuantization

Statistical Model Compression for Small-Footprint Natural Language Understanding

2018-07-19 · Grant P. Strimel, Kanthashree Mysore Sathyendra, Stanislav Peshterliev

In this paper we investigate statistical model compression applied to natural language understanding (NLU) models. Small-footprint NLU models are important for enabling offline systems on hardware restricted devices, and…

Model CompressionNatural Language UnderstandingQuantization

Device Tuning for Multi-Task Large Model

2023-02-21 · Penghao Jiang, Xuanchen Hou, Yinsi Zhou

Unsupervised pre-training approaches have achieved great success in many fields such as Computer Vision (CV), Natural Language Processing (NLP) and so on. However, compared to typical deep learning models, pre-training o…

modelMulti-Task LearningUnsupervised Pre-training

Extreme Compression of Large Language Models via Additive Quantization

2024-01-11 · Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar 외

The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices. In this paper, we revisit the problem of "ext…

CPUGPUInformation RetrievalQuantization

Map-Assisted Remote-Sensing Image Compression at Extremely Low Bitrates

2024-09-03 · Yixuan Ye, Ce Wang, Wanjie Sun, Zhenzhong Chen

Remote-sensing (RS) image compression at extremely low bitrates has always been a challenging task in practical scenarios like edge device storage and narrow bandwidth transmission. Generative models including VAEs and G…

Image Compression