Category Enhanced Word Embedding
Distributed word representations have been demonstrated to be effective in capturing semantic and syntactic regularities. Unsupervised representation learning from large unlabeled corpora can learn similar representations for those words that present similar co-occurrence statistics. Besides local occurrence statistics, global topical information is also important knowledge that may help discriminate a word from another. In this paper, we incorporate category information of documents in the learning of word representations and to learn the proposed models in a document-wise manner. Our models outperform several state-of-the-art models in word analogy and word similarity tasks. Moreover, we evaluate the learned word vectors on sentiment analysis and text classification tasks, which shows the superiority of our learned word vectors. We also learn high-quality category embeddings that reflect topical meanings.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationRepresentation LearningSentiment Analysistext-classificationText ClassificationWord SimilaritySimilar Papers 제목 키워드 기반
Encoding Category Trees Into Word-Embeddings Using Geometric Approach
We present a novel method to implicitly encode a tree-structured category information into word-embeddings, resulting in super-dimensional ball representations ($n$-ball embedding for short). Inclusion relations among $n…
Word EmbeddingsLearning Joint Embedding with Modality Alignments for Cross-Modal Retrieval of Recipes and Food Images
This paper presents a three-tier modality alignment approach to learning text-image joint embedding, coined as JEMA, for cross-modal retrieval of cooking recipes and food images. The first tier improves recipe text embed…
cross-modal alignmentCross-Modal RetrievalRetrievalTerm Extraction+1Learning Text-Image Joint Embedding for Efficient Cross-Modal Retrieval with Deep Feature Engineering
This paper introduces a two-phase deep feature engineering framework for efficient learning of semantics enhanced joint embedding, which clearly separates the deep feature engineering in data preprocessing from training …
Cross-Modal RetrievalFeature EngineeringRetrievalTripletDeepCAT: Deep Category Representation for Query Understanding in E-commerce Search
Mapping a search query to a set of relevant categories in the product taxonomy is a significant challenge in e-commerce search for two reasons: 1) Training data exhibits severe class imbalance problem due to biased click…
Exploring Semantic Spaces for Detecting Clustering and Switching in Verbal Fluency
In this work, we explore the fitness of various word/concept representations in analyzing an experimental verbal fluency dataset providing human responses to 10 different category enumeration tasks. Based on human annota…
Clustering