paper-with-me

홈 › Papers

Leveraging web resources for keyword assignment to short text documents

2017-06-19 · Singhal Ayush, Kasturi Ravindra, Sharma Ankit, Srivastava Jaideep

Assigning relevant keywords to documents is very important for efficient retrieval, clustering and management of the documents. Especially with the web corpus deluged with digital documents, automation of this task is of prime importance. Keyword assignment is a broad topic of research which refers to tagging of document with keywords, key-phrases or topics. For text documents, the keyword assignment techniques have been developed under two sub-topics: automatic keyword extraction (AKE) and automatic key-phrase abstraction. However, the approaches developed in the literature for full text documents cannot be used to assign keywords to low text content documents like twitter feeds, news clips, product reviews or even short scholarly text. In this work, we point out several practical challenges encountered in tagging such low text content documents. As a solution to these challenges, we show that the proposed approaches which leverage knowledge from several open source web resources enhance the quality of the tags (keywords) assigned to the low text content documents. The performance of the proposed approach is tested on real world corpus consisting of scholarly documents with text content ranging from only the text in the title of the document (5-10 words) to the summary text/abstract (100- 150 words). We find that the proposed approach not just improves the accuracy of keyword assignment but offer a computationally efficient solution which can be used in real world applications.

📄 PDF Abstract BibTeX arXiv:1706.05985

Code (0)

등록된 구현이 없습니다.

Tasks

Keyword ExtractionManagementRetrieval

Similar Papers 제목 키워드 기반

S2vNTM: Semi-supervised vMF Neural Topic Modeling

2023-07-06 · Weijie Xu, Jay Desai, Srinivasan Sengamedu, Xiaoyu Jiang 외

Language model based methods are powerful techniques for text classification. However, the models have several shortcomings. (1) It is difficult to integrate human knowledge such as keywords. (2) It needs a lot of resour…

Language ModelingLanguage Modellingtext-classificationText Classification

Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings

2025-08-18 · Taous Iatariene, Alexandre Guérin, Romain Serizel arxiv

Speaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common…

Knowledge Distillation

Geological Inference from Textual Data using Word Embeddings

2025-04-10 · Nanmanas Linphrachaya, Irving Gómez-Méndez, Adil Siripatana

This research explores the use of Natural Language Processing (NLP) techniques to locate geological resources, with a specific focus on industrial minerals. By using word embeddings trained with the GloVe model, we extra…

BenchmarkingWord Embeddings

A critical survey on measuring success in rank-based keyword assignment to documents

2015-06-01 · JEPTALNRECITAL 2015 6 · Natalie Schluter

Evaluation approaches for unsupervised rank-based keyword assignment are nearly as numerous as are the existing systems. The prolific production of each newly used metric (or metric twist) seems to stem from general dis-…

Topic Modeling based on Keywords and Context

2017-10-07 · Johannes Schneider

Current topic models often suffer from discovering topics not matching human intuition, unnatural switching of topics within documents and high computational demands. We address these concerns by proposing a topic model …

General ClassificationTopic Models