paper-with-me

홈 › Papers

Study of keyword extraction techniques for Electric Double Layer Capacitor domain using text similarity indexes: An experimental analysis

2021-11-13 · M. Saef Ullah Miah, Junaida Sulaiman, Talha Bin Sarwar, Kamal Z. Zamli, Rajan Jose

Keywords perform a significant role in selecting various topic-related documents quite easily. Topics or keywords assigned by humans or experts provide accurate information. However, this practice is quite expensive in terms of resources and time management. Hence, it is more satisfying to utilize automated keyword extraction techniques. Nevertheless, before beginning the automated process, it is necessary to check and confirm how similar expert-provided and algorithm-generated keywords are. This paper presents an experimental analysis of similarity scores of keywords generated by different supervised and unsupervised automated keyword extraction algorithms with expert provided keywords from the Electric Double Layer Capacitor (EDLC) domain. The paper also analyses which texts provide better keywords like positive sentences or all sentences of the document. From the unsupervised algorithms, YAKE, TopicRank, MultipartiteRank, and KPMiner are employed for keyword extraction. From the supervised algorithms, KEA and WINGNUS are employed for keyword extraction. To assess the similarity of the extracted keywords with expert-provided keywords, Jaccard, Cosine, and Cosine with word vector similarity indexes are employed in this study. The experiment shows that the MultipartiteRank keyword extraction technique measured with cosine with word vector similarity index produces the best result with 92% similarity with expert provided keywords. This study can help the NLP researchers working with the EDLC domain or recommender systems to select more suitable keyword extraction and similarity index calculation techniques.

📄 PDF Abstract BibTeX arXiv:2111.07068

Code (1)

ping543f/kwd-extraction-study 공식 구현

Tasks

Keyword ExtractionManagementRecommendation Systemstext similarity

Methods 이 논문이 사용한 방법론

Electric Electric is an energy-based cloze model for representation learning over text. Like BERT, it is a conditional generative model of tokens given their contexts. However,…

Similar Papers 제목 키워드 기반

Exploring Linguistically-Lightweight Keyword Extraction Techniques for Indexing News Articles in a Multilingual Set-up

2021-04-01 · EACL (Hackashop) 2021 4 · Jakub Piskorski, Nicolas Stefanovitch, Guillaume Jacquet, Aldo Podavini

This paper presents a study of state-of-the-art unsupervised and linguistically unsophisticated keyword extraction algorithms, based on statistic-, graph-, and embedding-based approaches, including, i.a., Total Keyword F…

ArticlesKeyword Extraction

Comparative Study of Domain Driven Terms Extraction Using Large Language Models

2024-04-02 · Sandeep Chataut, Tuyen Do, Bichar Dip Shrestha Gurung, Shiva Aryal 외

Keywords play a crucial role in bridging the gap between human understanding and machine processing of textual data. They are essential to data enrichment because they form the basis for detailed annotations that provide…

Document SummarizationHallucinationInformation RetrievalKeyword Extraction+2

Enhancing Retrieval-Augmented Generation for Electric Power Industry Customer Support

2025-08-01 · Hei Yu Chan, Kuok Tou Ho, Chenglong Ma, Yujing Si 외 arxiv

Many AI customer service systems use standard NLP pipelines or finetuned language models, which often fall short on ambiguous, multi-intent, or detail-specific queries. This case study evaluates recent techniques: query …

Intent Recognition

MAKED: Multi-lingual Automatic Keyword Extraction Dataset

2022-06-01 · LREC 2022 6 · Yash Verma, Anubhav Jangra, Sriparna Saha, Adam Jatowt 외

Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification. Development and evaluation of keyword extraction techniques require an exhaustive dataset; h…

ArticlesClusteringDiversityKeyword Extraction

Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI

2026-07-10 · Miguel Arana-Catania, Catherine Conisbee, Matthew Kidd arxiv

Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections. This article reports the findings of the "Extracting Keywords from Crowdsourced Collections" proj…

Keyword Extraction