Toward Selectivity Based Keyword Extraction for Croatian News
Preliminary report on network based keyword extraction for Croatian is an unsupervised method for keyword extraction from the complex network. We build our approach with a new network measure the node selectivity, motivated by the research of the graph based centrality approaches. The node selectivity is defined as the average weight distribution on the links of the single node. We extract nodes (keyword candidates) based on the selectivity value. Furthermore, we expand extracted nodes to word-tuples ranked with the highest in/out selectivity values. Selectivity based extraction does not require linguistic knowledge while it is purely derived from statistical and structural information en-compassed in the source text which is reflected into the structure of the network. Obtained sets are evaluated on a manually annotated keywords: for the set of extracted keyword candidates average F1 score is 24,63%, and average F2 score is 21,19%; for the exacted words-tuples candidates average F1 score is 25,9% and average F2 score is 24,47%.
Code (0)
등록된 구현이 없습니다.
Tasks
Keyword ExtractionSimilar Papers 제목 키워드 기반
Extending Neural Keyword Extraction with TF-IDF tagset matching
Keyword extraction is the task of identifying words (or multi-word expressions) that best describe a given document and serve in news portals to link articles of similar topics. In this work we develop and evaluate our m…
ArticlesKeyword ExtractionOut of Thin Air: Is Zero-Shot Cross-Lingual Keyword Detection Better Than Unsupervised?
Keyword extraction is the task of retrieving words that are essential to the content of a given document. Researchers proposed various approaches to tackle this problem. At the top-most level, approaches are divided into…
Keyword ExtractionPretrained Multilingual Language ModelsToward Network-based Keyword Extraction from Multitopic Web Documents
In this paper we analyse the selectivity measure calculated from the complex network in the task of the automatic keyword extraction. Texts, collected from different web sources (portals, forums), are represented as dire…
Keyword ExtractionSentenceQuotations, Coreference Resolution, and Sentiment Annotations in Croatian News Articles: An Exploratory Study
This paper presents a corpus annotated for the task of direct-speech extraction in Croatian. The paper focuses on the annotation of the quotation, co-reference resolution, and sentiment annotation in SETimes news corpus …
Articlescoreference-resolutionCoreference ResolutionSpeech ExtractionComplex Networks Measures for Differentiation between Normal and Shuffled Croatian Texts
This paper studies the properties of the Croatian texts via complex networks. We present network properties of normal and shuffled Croatian texts for different shuffling principles: on the sentence level and on the text …
Sentence