Keywords lie far from the mean of all words in local vector space
Keyword extraction is an important document process that aims at finding a small set of terms that concisely describe a document's topics. The most popular state-of-the-art unsupervised approaches belong to the family of the graph-based methods that build a graph-of-words and use various centrality measures to score the nodes (candidate keywords). In this work, we follow a different path to detect the keywords from a text document by modeling the main distribution of the document's words using local word vector representations. Then, we rank the candidates based on their position in the text and the distance between the corresponding local vectors and the main distribution's center. We confirm the high performance of our approach compared to strong baselines and state-of-the-art unsupervised keyword extraction methods, through an extended experimental study, investigating the properties of the local representations.
Code (1)
Tasks
AllKeyword ExtractionPositionSimilar Papers 제목 키워드 기반
A Generalized Vector Space Model for Ontology-Based Information Retrieval
Named entities (NE) are objects that are referred to by names such as people, organizations and locations. Named entities and keywords are important to the meaning of a document. We propose a generalized vector space mod…
Information RetrievalRetrievalExploring Combinations of Ontological Features and Keywords for Text Retrieval
Named entities have been considered and combined with keywords to enhance information retrieval performance. However, there is not yet a formal and complete model that takes into account entity names, classes, and identi…
Information RetrievalRetrievalText RetrievalInformation Retrieval in long documents: Word clustering approach for improving Semantics
In this paper, we propose an alternative to deep neural networks for semantic information retrieval for the case of long documents. This new approach exploiting clustering techniques to take into account the meaning of w…
ClusteringInformation RetrievalRetrievalSearching for Discriminative Words in Multidimensional Continuous Feature Space
Word feature vectors have been proven to improve many NLP tasks. With recent advances in unsupervised learning of these feature vectors, it became possible to train it with much more data, which also resulted in better q…
Part-Of-Speech TaggingRPM-Oriented Query Rewriting Framework for E-commerce Keyword-Based Sponsored Search
Sponsored search optimizes revenue and relevance, which is estimated by Revenue Per Mille (RPM). Existing sponsored search models are all based on traditional statistical models, which have poor RPM performance when quer…