Learning Only from Relevant Keywords and Unlabeled Documents
We consider a document classification problem where document labels are absent but only relevant keywords of a target class and unlabeled documents are given. Although heuristic methods based on pseudo-labeling have been considered, theoretical understanding of this problem has still been limited. Moreover, previous methods cannot easily incorporate well-developed techniques in supervised text classification. In this paper, we propose a theoretically guaranteed learning framework that is simple to implement and has flexible choices of models, e.g., linear models or neural networks. We demonstrate how to optimize the area under the receiver operating characteristic curve (AUC) effectively and also discuss how to adjust it to optimize other well-known evaluation metrics such as the accuracy and F1-measure. Finally, we show the effectiveness of our framework using benchmark datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Document ClassificationGeneral Classificationtext-classificationText ClassificationSimilar Papers 제목 키워드 기반
FastClass: A Time-Efficient Approach to Weakly-Supervised Text Classification
Weakly-supervised text classification aims to train a classifier using only class descriptions and unlabeled data. Recent research shows that keyword-driven methods can achieve state-of-the-art performance on various tas…
Classificationtext-classificationText ClassificationWeakly Supervised ClassificationLbl2Vec: An Embedding-Based Approach for Unsupervised Document Retrieval on Predefined Topics
In this paper, we consider the task of retrieving documents with predefined topics from an unlabeled document dataset using an unsupervised approach. The proposed unsupervised approach requires only a small number of key…
Document ClassificationRetrievalUnsupervised Text ClassificationWorld KnowledgeLeveraging web resources for keyword assignment to short text documents
Assigning relevant keywords to documents is very important for efficient retrieval, clustering and management of the documents. Especially with the web corpus deluged with digital documents, automation of this task is of…
Keyword ExtractionManagementRetrievalPositive unlabeled learning for building recommender systems in a parliamentary setting
Our goal is to learn about the political interests and preferences of the Members of Parliament by mining their parliamentary activity, in order to develop a recommendation/filtering system that, given a stream of docume…
Information RetrievalRecommendation SystemsRetrievalTopicSifter: Interactive Search Space Reduction Through Targeted Topic Modeling
Topic modeling is commonly used to analyze and understand large document collections. However, in practice, users want to focus on specific aspects or "targets" rather than the entire corpus. For example, given a large c…
Retrieval