paper-with-me

홈 › Papers

On the Downstream Performance of Compressed Word Embeddings

2019-09-03 · NeurIPS 2019 12 · Avner May, Jian Zhang, Tri Dao, Christopher Ré

Compressing word embeddings is important for deploying NLP models in memory-constrained settings. However, understanding what makes compressed embeddings perform well on downstream tasks is challenging---existing measures of compression quality often fail to distinguish between embeddings that perform well and those that do not. We thus propose the eigenspace overlap score as a new measure. We relate the eigenspace overlap score to downstream performance by developing generalization bounds for the compressed embeddings in terms of this score, in the context of linear and logistic regression. We then show that we can lower bound the eigenspace overlap score for a simple uniform quantization compression method, helping to explain the strong empirical performance of this method. Finally, we show that by using the eigenspace overlap score as a selection criterion between embeddings drawn from a representative set we compressed, we can efficiently identify the better performing embedding with up to $2\times$ lower selection error rates than the next best measure of compression quality, and avoid the cost of training a model for each task of interest.

📄 PDF Abstract BibTeX arXiv:1909.01264

Code (1)

HazyResearch/smallfry 공식 구현 pytorch

Tasks

Generalization BoundsQuantizationWord Embeddings

Similar Papers 제목 키워드 기반

Adaptive Compression of Word Embeddings

2020-07-01 · ACL 2020 6 · Yeachan Kim, Kang-Min Kim, SangKeun Lee

Distributed representations of words have been an indispensable component for natural language processing (NLP) tasks. However, the large memory footprint of word embeddings makes it challenging to deploy NLP models to m…

Self-Driving CarsWord Embeddings

A Compressed Sensing View of Unsupervised Text Embeddings, Bag-of-n-Grams, and LSTMs

2018-01-01 · ICLR 2018 1 · Sanjeev Arora, Mikhail Khodak, Nikunj Saunshi, Kiran Vodrahalli

Low-dimensional vector embeddings, computed using LSTMs or simpler techniques, are a popular approach for capturing the “meaning” of text and a form of unsupervised learning useful for downstream tasks. However, their po…

compressed sensing

An Empirical Study of the Downstream Reliability of Pre-Trained Word Embeddings

2020-12-01 · COLING 2020 8 · Anthony Rios, Brandon Lwowski

While pre-trained word embeddings have been shown to improve the performance of downstream tasks, many questions remain regarding their reliability: Do the same pre-trained word embeddings result in the best performance …

ImputationWord Embeddings

Direction is what you need: Improving Word Embedding Compression in Large Language Models

2021-06-15 · ACL (RepL4NLP) 2021 8 · Klaudia Bałazy, Mohammadreza Banaei, Rémi Lebret, Jacek Tabor 외

The adoption of Transformer-based models in natural language processing (NLP) has led to great success using a massive number of parameters. However, due to deployment constraints in edge devices, there has been a rising…

Language ModelingLanguage Modelling

Reconstructing Word Embeddings via Scattered $k$-Sub-Embedding

2021-09-29 · Soonyong Hwang, Byung-Ro Moon

The performance of modern neural language models relies heavily on the diversity of the vocabularies. Unfortunately, the language models tend to cover more vocabularies, the embedding parameters in the language models su…

DiversityWord Embeddings