Text Relatedness Based on a Word Thesaurus
The computation of relatedness between two fragments of text in an automated manner requires taking into account a wide range of factors pertaining to the meaning the two fragments convey, and the pairwise relations between their words. Without doubt, a measure of relatedness between text segments must take into account both the lexical and the semantic relatedness between words. Such a measure that captures well both aspects of text relatedness may help in many tasks, such as text retrieval, classification and clustering. In this paper we present a new approach for measuring the semantic relatedness between words based on their implicit semantic links. The approach exploits only a word thesaurus in order to devise implicit semantic links between words. Based on this approach, we introduce Omiotis, a new measure of semantic relatedness between texts which capitalizes on the word-to-word semantic relatedness measure (SR) and extends it to measure the relatedness between texts. We gradually validate our method: we first evaluate the performance of the semantic relatedness measure between individual words, covering word-to-word similarity and relatedness, synonym identification and word analogy; then, we proceed with evaluating the performance of our method in measuring text-to-text semantic relatedness in two tasks, namely sentence-to-sentence similarity and paraphrase recognition. Experimental evaluation shows that the proposed method outperforms every lexicon-based method of semantic relatedness in the selected tasks and the used data sets, and competes well against corpus-based and hybrid approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringRetrievalSentenceSentence SimilarityText RetrievalWord SimilaritySimilar Papers 제목 키워드 기반
Can Network Embedding of Distributional Thesaurus be Combined with Word Vectors for Better Representation?
Distributed representations of words learned from text have proved to be successful in various natural language processing tasks in recent times. While some methods represent words as vectors computed from text using pre…
Network EmbeddingWord SimilarityHuman and Machine Judgements for Russian Semantic Relatedness
Semantic relatedness of terms represents similarity of meaning by a numerical score. On the one hand, humans easily make judgments about semantic relatedness. On the other hand, this kind of information is useful in lang…
HG2Vec: Improved Word Embeddings from Dictionary and Thesaurus Based Heterogeneous Graph
Learning word embeddings is an essential topic in natural language processing. Most existing works use a vast corpus as a primary source while training, but this requires massive time and space for data pre-processing an…
Learning Word EmbeddingsWord EmbeddingsWord SimilaritySRHR at SemEval-2017 Task 6: Word Associations for Humour Recognition
This paper explores the role of semantic relatedness features, such as word associations, in humour recognition. Specifically, we examine the task of inferring pairwise humour judgments in Twitter hashtag wars. We examin…
AvgUsing Thesaurus Data to Improve Coreference Resolution for Russian
Semantic information about entities, specifically, how close in meaning two mentions are to each other, can become very useful for the task of co-reference resolution. One of the most well-researched and widely used form…
coreference-resolutionCoreference ResolutionSemantic SimilaritySemantic Textual Similarity