ELSKE: Efficient Large-Scale Keyphrase Extraction
Keyphrase extraction methods can provide insights into large collections of documents such as social media posts. Existing methods, however, are less suited for the real-time analysis of streaming data, because they are computationally too expensive or require restrictive constraints regarding the structure of keyphrases. We propose an efficient approach to extract keyphrases from large document collections and show that the method also performs competitively on individual documents.
Code (1)
Tasks
Information RetrievalKeyphrase ExtractionKeyword ExtractionMulti-Document SummarizationText SummarizationSimilar Papers 제목 키워드 기반
Theme-driven Keyphrase Extraction to Analyze Social Media Discourse
Social media platforms are vital resources for sharing self-reported health experiences, offering rich data on various health topics. Despite advancements in Natural Language Processing (NLP) enabling large-scale social …
Keyphrase ExtractionLanguage ModellingLarge Language ModelLarge-Scale Evaluation of Keyphrase Extraction Models
Keyphrase extraction models are usually evaluated under different, not directly comparable, experimental setups. As a result, it remains unclear how well proposed models actually perform, and how they compare to each oth…
Keyphrase ExtractionOpen Domain Web Keyphrase Extraction Beyond Language Modeling
This paper studies keyphrase extraction in real-world scenarios where documents are from diverse domains and have variant content quality. We curate and release OpenKP, a large scale open domain keyphrase extraction data…
Keyphrase ExtractionLanguage ModelingLanguage ModellingLocal Word Vectors Guiding Keyphrase Extraction
Automated keyphrase extraction is a fundamental textual information processing task concerned with the selection of representative phrases from a document that summarize its content. This work presents a novel unsupervis…
Keyphrase ExtractionWord EmbeddingsSlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction
Keyphrase extraction for morphologically rich, low-resource languages remains understudied, largely due to the scarcity of suitable evaluation datasets. We address this gap for Slovak by constructing a dataset of 227,432…
Keyphrase Extraction