Incorporating Subword Information into Matrix Factorization Word Embeddings
The positive effect of adding subword information to word embeddings has been demonstrated for predictive models. In this paper we investigate whether similar benefits can also be derived from incorporating subwords into counting models. We evaluate the impact of different types of subwords (n-grams and unsupervised morphemes), with results confirming the importance of subword information in learning representations of rare and out-of-vocabulary words.
Code (1)
Tasks
Word EmbeddingsSimilar Papers 제목 키워드 기반
Guided Semi-Supervised Non-negative Matrix Factorization on Legal Documents
Classification and topic modeling are popular techniques in machine learning that extract information from large-scale datasets. By incorporating a priori information such as labels or important features, methods have be…
ClassificationIncorporating Side Information in Probabilistic Matrix Factorization with Gaussian Processes
Probabilistic matrix factorization (PMF) is a powerful method for modeling data associ- ated with pairwise relationships, Finding use in collaborative Filtering, computational bi- ology, and document analysis, among othe…
Collaborative FilteringGaussian ProcessesParVecMF: A Paragraph Vector-based Matrix Factorization Recommender System
Review-based recommender systems have gained noticeable ground in recent years. In addition to the rating scores, those systems are enriched with textual evaluations of items by the users. Neural language processing mode…
Recommendation SystemsLGLMF: Local Geographical based Logistic Matrix Factorization Model for POI Recommendation
With the rapid growth of Location-Based Social Networks, personalized Points of Interest (POIs) recommendation has become a critical task to help users explore their surroundings. Due to the scarcity of check-in data, th…
Community Detection in Political Twitter Networks using Nonnegative Matrix Factorization Methods
Community detection is a fundamental task in social network analysis. In this paper, first we develop an endorsement filtered user connectivity network by utilizing Heider's structural balance theory and certain Twitter …
ClusteringCommunity DetectionWord Similarity