paper-with-me

홈 › Papers

Neural Networks Compression for Language Modeling

2017-08-20 · Artem M. Grachev, Dmitry I. Ignatov, Andrey V. Savchenko

In this paper, we consider several compression techniques for the language modeling problem based on recurrent neural networks (RNNs). It is known that conventional RNNs, e.g, LSTM-based networks in language modeling, are characterized with either high space complexity or substantial inference time. This problem is especially crucial for mobile applications, in which the constant interaction with the remote server is inappropriate. By using the Penn Treebank (PTB) dataset we compare pruning, quantization, low-rank factorization, tensor train decomposition for LSTM networks in terms of model size and suitability for fast inference.

📄 PDF Abstract BibTeX arXiv:1708.05963

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingQuantization

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Lossless Data Compression with Transformer

2019-09-25 · Gautier Izacard, Armand Joulin, Edouard Grave

Transformers have replaced long-short term memory and other recurrent neural networks variants in sequence modeling. It achieves state-of-the-art performance on a wide range of tasks related to natural language processin…

Data CompressionLanguage ModelingLanguage ModellingMachine Translation+2

CCF: A Context Compression Framework for Efficient Long-Sequence Language Modeling

2025-09-11 · Wenhao Li, Bangcheng Sun, Weihao Ye, Tianyi Zhang 외 arxiv

Scaling language models to longer contexts is essential for capturing rich dependencies across extended discourse. However, naïve context extension imposes significant computational and memory burdens, often resulting in…

Compression of Recurrent Neural Networks for Efficient Language Modeling

2019-02-06 · Artem M. Grachev, Dmitry I. Ignatov, Andrey V. Savchenko

Recurrent neural networks have proved to be an effective method for statistical language modeling. However, in practice their memory and run-time complexity are usually too large to be implemented in real-time offline mo…

Language ModelingLanguage ModellingQuantization

Proxy Compression for Language Modeling

2026-02-04 · Lin Zheng, Xinyu Li, Qian Liu, Xiachong Feng 외 arxiv

Modern language models are trained almost exclusively on token sequences produced by a fixed tokenizer, an external lossless compressor often over UTF-8 byte sequences, thereby coupling the model to that compressor. This…

Exploring Effective Mask Sampling Modeling for Neural Image Compression

2023-06-09 · Lin Liu, Mingming Zhao, Shanxin Yuan, Wenlong Lyu 외

Image compression aims to reduce the information redundancy in images. Most existing neural image compression methods rely on side information from hyperprior or context models to eliminate spatial redundancy, but rarely…

Image CompressionSelf-Supervised Learning