paper-with-me

홈 › Papers

Global dense vector representations for words or items using shared parameter alternating Tweedie model

2024-12-31 · Taejoon Kim, HaiYan Wang

In this article, we present a model for analyzing the cooccurrence count data derived from practical fields such as user-item or item-item data from online shopping platform, cooccurring word-word pairs in sequences of texts. Such data contain important information for developing recommender systems or studying relevance of items or words from non-numerical sources. Different from traditional regression models, there are no observations for covariates. Additionally, the cooccurrence matrix is typically of so high dimension that it does not fit into a computer's memory for modeling. We extract numerical data by defining windows of cooccurrence using weighted count on the continuous scale. Positive probability mass is allowed for zero observations. We present Shared parameter Alternating Tweedie (SA-Tweedie) model and an algorithm to estimate the parameters. We introduce a learning rate adjustment used along with the Fisher scoring method in the inner loop to help the algorithm stay on track of optimizing direction. Gradient descent with Adam update was also considered as an alternative method for the estimation. Simulation studies and an application showed that our algorithm with Fisher scoring and learning rate adjustment outperforms the other two methods. Pseudo-likelihood approach with alternating parameter update was also studied. Numerical studies showed that the pseudo-likelihood approach is not suitable in our shared parameter alternating regression models with unobserved covariates.

📄 PDF Abstract BibTeX arXiv:2501.00623

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation Systems

Methods 이 논문이 사용한 방법론

Adam 설명 없음

Similar Papers 제목 키워드 기반

HLP@UPenn at SemEval-2017 Task 4A: A simple, self-optimizing text classification system combining dense and sparse vectors

2017-08-01 · SEMEVAL 2017 8 · Abeed Sarker, Graciela Gonzalez

We present a simple supervised text classification system that combines sparse and dense vector representations of words, and generalized representations of words via clusters. The sparse vectors are generated from word …

ClassificationEpidemiologyGeneral ClassificationSentiment Analysis+2

Semantic Representations of Word Senses and Concepts

2016-08-02 · José Camacho-Collados, Ignacio Iacobacci, Roberto Navigli, Mohammad Taher Pilehvar

Representing the semantics of linguistic items in a machine-interpretable form has been a major goal of Natural Language Processing since its earliest days. Among the range of different linguistic items, words have attra…

Distributed Representations of Sentences and Documents

2014-05-16 · Quoc V. Le, Tomas Mikolov

Many machine learning algorithms require the input to be represented as a fixed-length feature vector. When it comes to texts, one of the most common fixed-length features is bag-of-words. Despite their popularity, bag-o…

Question AnsweringSentiment AnalysisText Classification

Sparse Lifting of Dense Vectors: Unifying Word and Sentence Representations

2019-11-05 · Wenye Li, Senyue Hao

As the first step in automated natural language processing, representing words and sentences is of central importance and has attracted significant research attention. Different approaches, from the early one-hot and bag…

Sentence

Semantic Vector Encoding and Similarity Search Using Fulltext Search Engines

2017-08-01 · WS 2017 8 · Jan Rygl, Jan Pomik{\'a}lek, Radim {\v{R}}eh{\r{u}}{\v{r}}ek, Michal R{\r{u}}{\v{z}}i{\v{c}}ka 외

Vector representations and vector space modeling (VSM) play a central role in modern machine learning. We propose a novel approach to {`}vector similarity searching{'} over dense semantic representations of words and doc…

Information RetrievalRepresentation Learning