Toward Network-based Keyword Extraction from Multitopic Web Documents
In this paper we analyse the selectivity measure calculated from the complex network in the task of the automatic keyword extraction. Texts, collected from different web sources (portals, forums), are represented as directed and weighted co-occurrence complex networks of words. Words are nodes and links are established between two nodes if they are directly co-occurring within the sentence. We test different centrality measures for ranking nodes - keyword candidates. The promising results are achieved using the selectivity measure. Then we propose an approach which enables extracting word pairs according to the values of the in/out selectivity and weight measures combined with filtering.
Code (0)
등록된 구현이 없습니다.
Tasks
Keyword ExtractionSentenceSimilar Papers 제목 키워드 기반
Topic Segmentation of Research Article Collections
Collections of research article data harvested from the web have become common recently since they are important resources for experimenting on tasks such as named entity recognition, text summarization, or keyword gener…
Articlesnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2Theme-weighted Ranking of Keywords from Text Documents using Phrase Embeddings
Keyword extraction is a fundamental task in natural language processing that facilitates mapping of documents to a concise set of representative single and multi-word phrases. Keywords from text documents are primarily e…
Keyword ExtractionKeyword Extraction in Scientific Documents
The scientific publication output grows exponentially. Therefore, it is increasingly challenging to keep track of trends and changes. Understanding scientific documents is an important step in downstream tasks such as kn…
Keyphrase ExtractionKeyword ExtractionLeveraging web resources for keyword assignment to short text documents
Assigning relevant keywords to documents is very important for efficient retrieval, clustering and management of the documents. Especially with the web corpus deluged with digital documents, automation of this task is of…
Keyword ExtractionManagementRetrieval