paper-with-me

홈 › Papers

Statistical Uncertainty in Word Embeddings: GloVe-V

2024-06-18 · Andrea Vallebueno, Cassandra Handan-Nader, Christopher D. Manning, Daniel E. Ho

Static word embeddings are ubiquitous in computational social science applications and contribute to practical decision-making in a variety of fields including law and healthcare. However, assessing the statistical uncertainty in downstream conclusions drawn from word embedding statistics has remained challenging. When using only point estimates for embeddings, researchers have no streamlined way of assessing the degree to which their model selection criteria or scientific conclusions are subject to noise due to sparsity in the underlying data used to generate the embeddings. We introduce a method to obtain approximate, easy-to-use, and scalable reconstruction error variance estimates for GloVe (Pennington et al., 2014), one of the most widely used word embedding models, using an analytical approximation to a multivariate normal model. To demonstrate the value of embeddings with variance (GloVe-V), we illustrate how our approach enables principled hypothesis testing in core word embedding tasks, such as comparing the similarity between different word pairs in vector space, assessing the performance of different models, and analyzing the relative degree of ethnic or gender bias in a corpus using different word lists.

📄 PDF Abstract BibTeX arXiv:2406.12165

Code (1)

reglab/glove-v 공식 구현

Tasks

Decision MakingModel SelectionWord Embeddings

Methods 이 논문이 사용한 방법론

GloVe GloVe Embeddings are a type of word embedding that encode the co-occurrence probability ratio between two words as vector differences. GloVe uses a weighted least squares…

Similar Papers 제목 키워드 기반

SemGloVe: Semantic Co-occurrences for GloVe from BERT

2020-12-30 · Leilei Gan, Zhiyang Teng, Yue Zhang, Linchao Zhu 외

GloVe learns word embeddings by leveraging statistical information from word co-occurrence matrices. However, word pairs in the matrices are extracted from a predefined local context window, which might lead to limited w…

Language ModelingLanguage ModellingWord EmbeddingsWord Similarity

A Bayesian approach to uncertainty in word embedding bias estimation

2023-06-15 · Alicja Dobrzeniecka, Rafal Urbaniak

Multiple measures, such as WEAT or MAC, attempt to quantify the magnitude of bias present in word embeddings in terms of a single-number metric. However, such metrics and the related statistical significance calculations…

Word Embeddings

WOVe: Incorporating Word Order in GloVe Word Embeddings

2021-05-18 · Mohammed Ibrahim, Susan Gauch, Tyler Gerth, Brandon Cox

Word vector representations open up new opportunities to extract useful information from unstructured text. Defining a word as a vector made it easy for the machine learning algorithms to understand a text and extract in…

Word EmbeddingsWord Similarity

Non-Linear Relational Information Probing in Word Embeddings

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Pre-trained word embeddings such as SkipGram and GloVe are known to contain a myriad of useful information about words. In this work, we use multilayer perceptrons (MLP) to probe the relational information contained in t…

RelationWord Embeddings

TeamDL at SemEval-2018 Task 8: Cybersecurity Text Analysis using Convolutional Neural Network and Conditional Random Fields

2018-06-01 · SEMEVAL 2018 6 · Manik R, an, Krishna Madgula, Snehanshu Saha

In this work we present our participation to SemEval-2018 Task 8 subtasks 1 {\&} 2 respectively. We developed Convolution Neural Network system for malware sentence classification (subtask 1) and Conditional Random Field…

General ClassificationLEMMAPOSSentence+4