paper-with-me

홈 › Papers

Learning Only from Relevant Keywords and Unlabeled Documents

2019-10-10 · IJCNLP 2019 11 · Nontawat Charoenphakdee, Jongyeong Lee, Yiping Jin, Dittaya Wanvarie, Masashi Sugiyama

We consider a document classification problem where document labels are absent but only relevant keywords of a target class and unlabeled documents are given. Although heuristic methods based on pseudo-labeling have been considered, theoretical understanding of this problem has still been limited. Moreover, previous methods cannot easily incorporate well-developed techniques in supervised text classification. In this paper, we propose a theoretically guaranteed learning framework that is simple to implement and has flexible choices of models, e.g., linear models or neural networks. We demonstrate how to optimize the area under the receiver operating characteristic curve (AUC) effectively and also discuss how to adjust it to optimize other well-known evaluation metrics such as the accuracy and F1-measure. Finally, we show the effectiveness of our framework using benchmark datasets.

📄 PDF Abstract BibTeX arXiv:1910.04385

Code (0)

등록된 구현이 없습니다.

Tasks

Document ClassificationGeneral Classificationtext-classificationText Classification

Similar Papers 제목 키워드 기반

FastClass: A Time-Efficient Approach to Weakly-Supervised Text Classification

2022-12-11 · Tingyu Xia, Yue Wang, Yuan Tian, Yi Chang

Weakly-supervised text classification aims to train a classifier using only class descriptions and unlabeled data. Recent research shows that keyword-driven methods can achieve state-of-the-art performance on various tas…

Classificationtext-classificationText ClassificationWeakly Supervised Classification

Lbl2Vec: An Embedding-Based Approach for Unsupervised Document Retrieval on Predefined Topics

2022-10-12 · Tim Schopf, Daniel Braun, Florian Matthes

In this paper, we consider the task of retrieving documents with predefined topics from an unlabeled document dataset using an unsupervised approach. The proposed unsupervised approach requires only a small number of key…

Document ClassificationRetrievalUnsupervised Text ClassificationWorld Knowledge

Leveraging web resources for keyword assignment to short text documents

2017-06-19 · Singhal Ayush, Kasturi Ravindra, Sharma Ankit, Srivastava Jaideep

Assigning relevant keywords to documents is very important for efficient retrieval, clustering and management of the documents. Especially with the web corpus deluged with digital documents, automation of this task is of…

Keyword ExtractionManagementRetrieval

Positive unlabeled learning for building recommender systems in a parliamentary setting

2024-01-19 · Luis M. de Camposa, Juan M. Fernández-Luna, Juan F. Huete, Luis Redondo-Expósito

Our goal is to learn about the political interests and preferences of the Members of Parliament by mining their parliamentary activity, in order to develop a recommendation/filtering system that, given a stream of docume…

Information RetrievalRecommendation SystemsRetrieval

TopicSifter: Interactive Search Space Reduction Through Targeted Topic Modeling

2019-07-28 · Hannah Kim, Dongjin Choi, Barry Drake, Alex Endert 외

Topic modeling is commonly used to analyze and understand large document collections. However, in practice, users want to focus on specific aspects or "targets" rather than the entire corpus. For example, given a large c…

Retrieval