Sentence Analogies: Exploring Linguistic Relationships and Regularities in Sentence Embeddings
While important properties of word vector representations have been studied extensively, far less is known about the properties of sentence vector representations. Word vectors are often evaluated by assessing to what degree they exhibit regularities with regard to relationships of the sort considered in word analogies. In this paper, we investigate to what extent commonly used sentence vector representation spaces as well reflect certain kinds of regularities. We propose a number of schemes to induce evaluation data, based on lexical analogy data as well as semantic relationships between sentences. Our experiments consider a wide range of sentence embedding methods, including ones based on BERT-style contextual embeddings. We find that different models differ substantially in their ability to reflect such regularities.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceSentence EmbeddingSentence-EmbeddingSentence EmbeddingsSimilar Papers 제목 키워드 기반
Sentence Analogies: Linguistic Regularities in Sentence Embeddings
While important properties of word vector representations have been studied extensively, far less is known about the properties of sentence vector representations. Word vectors are often evaluated by assessing to what de…
SentenceSentence EmbeddingSentence-EmbeddingSentence EmbeddingsInsights into Analogy Completion from the Biomedical Domain
Analogy completion has been a popular task in recent years for evaluating the semantic properties of word embeddings, but the standard methodology makes a number of assumptions about analogies that do not always hold, ei…
Word EmbeddingsAnalogies minus analogy test: measuring regularities in word embeddings
Vector space models of words have long been claimed to capture linguistic regularities as simple vector translations, but problems have been raised with this claim. We decompose and empirically analyze the classic arithm…
Word EmbeddingsParaphrases do not explain word analogies
Many types of distributional word embeddings (weakly) encode linguistic regularities as directions (the difference between "jump" and "jumped" will be in a similar direction to that of "walk" and "walked," and so on). Se…
Word EmbeddingsLong-form analogies generated by chatGPT lack human-like psycholinguistic properties
Psycholinguistic analyses provide a means of evaluating large language model (LLM) output and making systematic comparisons to human-generated text. These methods can be used to characterize the psycholinguistic properti…
FormLanguage ModelingLanguage ModellingLarge Language Model