paper-with-me

홈 › Papers

Estimator Vectors: OOV Word Embeddings based on Subword and Context Clue Estimates

2019-10-18 · Raj Patel, Carlotta Domeniconi

Semantic representations of words have been successfully extracted from unlabeled corpuses using neural network models like word2vec. These representations are generally high quality and are computationally inexpensive to train, making them popular. However, these approaches generally fail to approximate out of vocabulary (OOV) words, a task humans can do quite easily, using word roots and context clues. This paper proposes a neural network model that learns high quality word representations, subword representations, and context clue representations jointly. Learning all three types of representations together enhances the learning of each, leading to enriched word vectors, along with strong estimates for OOV words, via the combination of the corresponding context clue and subword embeddings. Our model, called Estimator Vectors (EV), learns strong word embeddings and is competitive with state of the art methods for OOV estimation.

📄 PDF Abstract BibTeX arXiv:1910.10491

Code (0)

등록된 구현이 없습니다.

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

PBoS: Probabilistic Bag-of-Subwords for Generalizing Word Embedding

2020-10-21 · Findings of the Association for Computational Linguistics 2020 · Zhao Jinman, Shawn Zhong, Xiaomin Zhang, YIngyu Liang

We look into the task of \emph{generalizing} word embeddings: given a set of pre-trained word vectors over a finite vocabulary, the goal is to predict embedding vectors for out-of-vocabulary words, \emph{without} extra c…

POSPOS TaggingWord EmbeddingsWord Similarity

Crossword: Estimating Unknown Embeddings using Cross Attention and Alignment Strategies

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Word embedding methods like word2vec and GloVe have been shown to learn strong representations of words. However, these methods only learn representations for words in the training corpus. This is problematic, as model…

Word Embeddings

Generalizing Word Embeddings using Bag of Subwords

2018-09-12 · EMNLP 2018 10 · Jinman Zhao, Sidharth Mudgal, YIngyu Liang

We approach the problem of generalizing pre-trained word embeddings beyond fixed-size vocabularies without using additional contextual information. We propose a subword-level word vector generation model that views words…

TAGWord EmbeddingsWord Similarity

Entropy-Based Subword Mining with an Application to Word Embeddings

2018-06-01 · WS 2018 6 · Ahmed El-Kishky, Frank Xu, Aston Zhang, Stephen Macke 외

Recent literature has shown a wide variety of benefits to mapping traditional one-hot representations of words and phrases to lower-dimensional real-valued vectors known as word embeddings. Traditionally, most word embed…

Language ModelingLanguage ModellingMachine TranslationSentiment Analysis+2

Treat the Word As a Whole or Look Inside? Subword Embeddings Model Language Change and Typology

2019-08-01 · WS 2019 8 · Yang Xu, Jiasheng Zhang, David Reitter

We use a variant of word embedding model that incorporates subword information to characterize the degree of compositionality in lexical semantics. Our models reveal some interesting yet contrastive patterns of long-term…