paper-with-me

홈 › Papers

Statistical Model Compression for Small-Footprint Natural Language Understanding

2018-07-19 · Grant P. Strimel, Kanthashree Mysore Sathyendra, Stanislav Peshterliev

In this paper we investigate statistical model compression applied to natural language understanding (NLU) models. Small-footprint NLU models are important for enabling offline systems on hardware restricted devices, and for decreasing on-demand model loading latency in cloud-based systems. To compress NLU models, we present two main techniques, parameter quantization and perfect feature hashing. These techniques are complementary to existing model pruning strategies such as L1 regularization. We performed experiments on a large scale NLU system. The results show that our approach achieves 14-fold reduction in memory usage compared to the original models with minimal predictive performance impact.

📄 PDF Abstract BibTeX arXiv:1807.07520

Code (0)

등록된 구현이 없습니다.

Tasks

Model CompressionNatural Language UnderstandingQuantization

Similar Papers 제목 키워드 기반

Extremely Small BERT Models from Mixed-Vocabulary Training

2019-09-25 · EACL 2021 2 · Sanqiang Zhao, Raghav Gupta, Yang song, Denny Zhou

Pretrained language models like BERT have achieved good results on NLP tasks, but are impractical on resource-limited devices due to memory footprint. A large fraction of this footprint comes from the input embeddings wi…

Knowledge DistillationLanguage ModellingModel CompressionWord Embeddings

SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression

2025-11-24 · Santhosh G S, Saurav Prakash, Balaraman Ravindran arxiv

Large Language Models (LLMs) face a significant bottleneck during autoregressive inference due to the massive memory footprint of the Key-Value (KV) cache. Existing compression techniques like token eviction, quantizatio…

Personalized Speech recognition on mobile devices

2016-03-10 · Ian McGraw, Rohit Prabhavalkar, Raziel Alvarez, Montse Gonzalez Arenas 외

We describe a large vocabulary speech recognition system that is accurate, has low latency, and yet has a small enough memory and computational footprint to run faster than real-time on a Nexus 5 Android smartphone. We e…

DecoderLanguage ModelingLanguage Modellingspeech-recognition+1

Low-Rank Tensor Approximation of Weights in Large Language Models via Cosine Lanczos Bidiagonalization

2026-01-23 · A. El Ichi, K. Jbilou arxiv

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language tasks but suffer from extremely large memory footprints and computational costs. In this paper, we introduce a tensor…

Distilling Neural Networks for Greener and Faster Dependency Parsing

2020-06-01 · WS 2020 7 · Mark Anderson, Carlos Gómez-Rodríguez

The carbon footprint of natural language processing research has been increasing in recent years due to its reliance on large and inefficient neural network implementations. Distillation is a network compression techniqu…

CPUDependency ParsingGPU