Reconstructing Word Embeddings via Scattered $k$-Sub-Embedding
The performance of modern neural language models relies heavily on the diversity of the vocabularies. Unfortunately, the language models tend to cover more vocabularies, the embedding parameters in the language models such as multilingual models used to occupy more than a half of their entire learning parameters. To solve this problem, we aim to devise a novel embedding structure to lighten the network without considerably performance degradation. To reconstruct $N$ embedding vectors, we initialize $k$ bundles of $M (\ll N)$ $k$-sub-embeddings to apply Cartesian product. Furthermore, we assign $k$-sub-embedding using the contextual relationship between tokens from pretrained language models. We adjust our $k$-sub-embedding structure to masked language models to evaluate proposed structure on downstream tasks. Our experimental results show that over 99.9$\%+$ compressed sub-embeddings for the language models performed comparably with the original embedding structure on GLUE and XNLI benchmarks.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityWord EmbeddingsSimilar Papers 제목 키워드 기반
Subword-based Compact Reconstruction of Word Embeddings
The idea of subword-based word embeddings has been proposed in the literature, mainly for solving the out-of-vocabulary (OOV) word problem observed in standard word-based word embeddings. In this paper, we propose a meth…
Word EmbeddingsTiny Word Embeddings Using Globally Informed Reconstruction
We reduce the model size of pre-trained word embeddings by a factor of 200 while preserving its quality. Previous studies in this direction created a smaller word embedding model by reconstructing pre-trained word repres…
Word EmbeddingsWord SimilarityDeconstructing and reconstructing word embedding algorithms
Uncontextualized word embeddings are reliable feature representations of words used to obtain high quality results for various NLP applications. Given the historical success of word embeddings in NLP, we propose a retros…
Word EmbeddingsComparing Euclidean and Hyperbolic Embeddings on the WordNet Nouns Hypernymy Graph
Nickel and Kiela (2017) present a new method for embedding tree nodes in the Poincare ball, and suggest that these hyperbolic embeddings are far more effective than Euclidean embeddings at embedding nodes in large, hiera…
MoRTy: Unsupervised Learning of Task-specialized Word Embeddings by Autoencoding
Word embeddings have undoubtedly revolutionized NLP. However, pretrained embeddings do not always work for a specific task (or set of tasks), particularly in limited resource setups. We introduce a simple yet effective, …
Word Embeddings