Non-distributional Word Vector Representations
Data-driven representation learning for words is a technique of central importance in NLP. While indisputably useful as a source of features in downstream tasks, such vectors tend to consist of uninterpretable components whose relationship to the categories of traditional lexical semantic theories is tenuous at best. We present a method for constructing interpretable word vectors from hand-crafted linguistic resources like WordNet, FrameNet etc. These vectors are binary (i.e, contain only 0 and 1) and are 99.9% sparse. We analyze their performance on state-of-the-art evaluation methods for distributional models of word vectors and find they are competitive to standard distributional approaches.
Code (1)
Tasks
Representation LearningSimilar Papers 제목 키워드 기반
Can Network Embedding of Distributional Thesaurus be Combined with Word Vectors for Better Representation?
Distributed representations of words learned from text have proved to be successful in various natural language processing tasks in recent times. While some methods represent words as vectors computed from text using pre…
Network EmbeddingWord SimilaritySentence Entailment in Compositional Distributional Semantics
Distributional semantic models provide vector representations for words by gathering co-occurrence frequencies from corpora of text. Compositional distributional models extend these from words to phrases and sentences. I…
SentencePost-Specialisation: Retrofitting Vectors of Words Unseen in Lexical Resources
Word vector specialisation (also known as retrofitting) is a portable, light-weight approach to fine-tuning arbitrary distributional word vector spaces by injecting external knowledge from rich lexical resources such as …
Dialogue State TrackingText SimplificationWord SimilaritySemantic Word Clusters Using Signed Normalized Graph Cuts
Vector space representations of words capture many aspects of word similarity, but such methods tend to make vector spaces in which antonyms (as well as synonyms) are close to each other. We present a new signed spectral…
ClusteringWord SimilarityBuilding Sense Representations in Danish by Combining Word Embeddings with Lexical Resources
Our aim is to identify suitable sense representations for NLP in Danish. We investigate sense inventories that correlate with human interpretations of word meaning and ambiguity as typically described in dictionaries and…
SentenceWord EmbeddingsWord Sense Disambiguation