An Improved Method for Class-specific Keyword Extraction: A Case Study in the German Business Registry
The task of $\textit{keyword extraction}$ is often an important initial step in unsupervised information extraction, forming the basis for tasks such as topic modeling or document classification. While recent methods have proven to be quite effective in the extraction of keywords, the identification of $\textit{class-specific}$ keywords, or only those pertaining to a predefined class, remains challenging. In this work, we propose an improved method for class-specific keyword extraction, which builds upon the popular $\textbf{KeyBERT}$ library to identify only keywords related to a class described by $\textit{seed keywords}$. We test this method using a dataset of German business registry entries, where the goal is to classify each business according to an economic sector. Our results reveal that our method greatly improves upon previous approaches, setting a new standard for $\textit{class-specific}$ keyword extraction.
Code (1)
Tasks
Document ClassificationKeyword ExtractionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Complex Network based Supervised Keyword Extractor
In this paper, we present a supervised framework for automatic keyword extraction from single document. We model the text as complex network, and construct the feature set by extracting select node properties from it. Se…
Keyphrase ExtractionKeyword ExtractionKeyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content
Brand reputation in the banking sector is maintained through insightful analysis of customer opinion on code-mixed and multilingual content. Conventional NLP models misclassify or ignore code-mixed text, when mix with lo…
Keyword ExtractionNERImproving Performance of Automatic Keyword Extraction (AKE) Methods Using PoS-Tagging and Enhanced Semantic-Awareness
Automatic keyword extraction (AKE) has gained more importance with the increasing amount of digital textual data that modern computing systems process. It has various applications in information retrieval (IR) and natura…
Information RetrievalKeyword ExtractionPOSPOS Tagging+1Back to the Basics: A Quantitative Analysis of Statistical and Graph-Based Term Weighting Schemes for Keyword Extraction
Term weighting schemes are widely used in Natural Language Processing and Information Retrieval. In particular, term weighting is the basis for keyword extraction. However, there are relatively few evaluation studies tha…
Information RetrievalKeyword ExtractionRetrievalSpecificityKeyword Extraction for Improved Document Retrieval in Conversational Search
Recent research has shown that mixed-initiative conversational search, based on the interaction between users and computers to clarify and improve a query, provides enormous advantages. Nonetheless, incorporating additio…
Conversational SearchKeyword ExtractionRetrievalSentence