paper-with-me

홈 › Papers

Investigating the Frequency Distortion of Word Embeddings and Its Impact on Bias Metrics

2022-11-15 · Francisco Valentini, Juan Cruz Sosa, Diego Fernandez Slezak, Edgar Altszyler

Recent research has shown that static word embeddings can encode word frequency information. However, little has been studied about this phenomenon and its effects on downstream tasks. In the present work, we systematically study the association between frequency and semantic similarity in several static word embeddings. We find that Skip-gram, GloVe and FastText embeddings tend to produce higher semantic similarity between high-frequency words than between other frequency combinations. We show that the association between frequency and similarity also appears when words are randomly shuffled. This proves that the patterns found are not due to real semantic associations present in the texts, but are an artifact produced by the word embeddings. Finally, we provide an example of how word frequency can strongly impact the measurement of gender bias with embedding-based metrics. In particular, we carry out a controlled experiment that shows that biases can even change sign or reverse their order by manipulating word frequencies.

📄 PDF Abstract BibTeX arXiv:2211.08203

Code (1)

ftvalentini/embeddingsfrequency 공식 구현

Tasks

Semantic SimilaritySemantic Textual SimilarityWord Embeddings

Methods 이 논문이 사용한 방법론

fastText fastText embeddings exploit subword information to construct word embeddings. Representations are learnt of character $n$-grams, and words represented as the sum of the…
GloVe GloVe Embeddings are a type of word embedding that encode the co-occurrence probability ratio between two words as vector differences. GloVe uses a weighted least squares…

Similar Papers 제목 키워드 기반

Frequency-based Distortions in Contextualized Word Embeddings

2021-04-17 · Kaitlyn Zhou, Kawin Ethayarajh, Dan Jurafsky

How does word frequency in pre-training data affect the behavior of similarity metrics in contextualized BERT embeddings? Are there systematic ways in which some word relationships are exaggerated or understated? In this…

Semantic SimilaritySemantic Textual SimilarityWord Embeddings

Frequency-aware Dimension Selection for Static Word Embedding by Mixed Product Distance

2023-05-13 · Lingfeng Shen, Haiyun Jiang, Lemao Liu, Ying Chen

Static word embedding is still useful, particularly for context-unavailable tasks, because in the case of no context available, pre-trained language models often perform worse than static word embeddings. Although dimens…

Word Embeddings

Investigating the Stability of Concrete Nouns in Word Embeddings

2019-05-01 · WS 2019 5 · B{\'e}n{\'e}dicte Pierrejean, Ludovic Tanguy

We know that word embeddings trained using neural-based methods (such as word2vec SGNS) are sensitive to stability problems and that across two models trained using the exact same set of parameters, the nearest neighbors…

Word Embeddings

Measuring Similarity by Linguistic Features rather than Frequency

2022-06-01 · ISA (LREC) 2022 6 · Rodolfo Delmonte, Nicolò Busetto

In the use and creation of current Deep Learning Models the only number that is used for the overall computation is the frequency value associated with the current word form in the corpus, which is used to substitute it.…

Word Embeddings

Evaluating Bias In Dutch Word Embeddings

2020-10-31 · GeBNLP (COLING) 2020 12 · Rodrigo Alejandro Chávez Mulsa, Gerasimos Spanakis

Recent research in Natural Language Processing has revealed that word embeddings can encode social biases present in the training data which can affect minorities in real world applications. This paper explores the gende…

ClusteringSentenceSentence EmbeddingsWord Embeddings