paper-with-me

Papers Unsupervised Text Classification

“Unsupervised Text Classification” 태그가 달린 논문 14편 · 필터 해제

Dual Refinement Cycle Learning: Unsupervised Text Classification of Mamba and Community Detection on Text Attributed Graph

2025-12-08 · Hong Wang, Yinglong Zhang, Hanhan Guo, Xuewen Xia 외 arxiv

Pretrained language models offer strong text understanding capabilities but remain difficult to deploy in real-world text-attributed networks due to their heavy dependence on labeled data. Meanwhile, community detection …

Unsupervised Text ClassificationRepresentation LearningCommunity Detection

One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification

2025-10-13 · Jens Van Nooten, Andriy Kosar, Guy De Pauw, Walter Daelemans arxiv

Distance-based unsupervised text classification is a method within text classification that leverages the semantic similarity between a label and a text to determine label relevance. This method provides numerous benefit…

Unsupervised Text ClassificationMulti-Label Text ClassificationMulti-Label ClassificationInformation Retrieval

Shuffle & Divide: Contrastive Learning for Long Text

2023-04-19 · Joonseok Lee, Seongho Joe, Kyoungwon Park, Bogun Kim 외

We propose a self-supervised learning method for long text documents based on contrastive learning. A key to our method is Shuffle and Divide (SaD), a simple text augmentation algorithm that sets up a pretext task requir…

Contrastive LearningDocument EmbeddingSelf-Supervised LearningText Augmentation+3

Text classification in shipping industry using unsupervised models and Transformer based supervised models

2022-12-21 · Ying Xie, Dongping Song

Obtaining labelled data in a particular context could be expensive and time consuming. Although different algorithms, including unsupervised learning, semi-supervised learning, self-learning have been adopted, the perfor…

ClassificationSelf-Learningtext-classificationText Classification+2

Evaluating Unsupervised Text Classification: Zero-shot and Similarity-based Approaches

2022-11-29 · Tim Schopf, Daniel Braun, Florian Matthes

Text classification of unseen classes is a challenging Natural Language Processing task and is mainly attempted using two different types of approaches. Similarity-based approaches attempt to classify instances based on …

Classificationtext-classificationText ClassificationUnsupervised Text Classification+1

Lbl2Vec: An Embedding-Based Approach for Unsupervised Document Retrieval on Predefined Topics

2022-10-12 · Tim Schopf, Daniel Braun, Florian Matthes

In this paper, we consider the task of retrieving documents with predefined topics from an unlabeled document dataset using an unsupervised approach. The proposed unsupervised approach requires only a small number of key…

Document ClassificationRetrievalUnsupervised Text ClassificationWorld Knowledge

Lex2Sent: A bagging approach to unsupervised sentiment analysis

2022-09-26 · Kai-Robin Lange, Jonas Rieger, Carsten Jentsch

Unsupervised text classification, with its most common form being sentiment analysis, used to be performed by counting words in a text that were stored in a lexicon, which assigns each word to one class or as a neutral w…

ClassificationDecoderGPUSentiment Analysis+5

DocSCAN: Unsupervised Text Classification via Learning from Neighbors

2021-05-09 · KONVENS (WS) 2022 9 · Dominik Stammbach, Elliott Ash

We introduce DocSCAN, a completely unsupervised text classification approach using Semantic Clustering by Adopting Nearest-Neighbors (SCAN). For each document, we obtain semantically informative vectors from a large pre-…

ClassificationClusteringGeneral ClassificationLanguage Modeling+6

Exclusive Topic Modeling

2021-02-06 · Hao Lei, Ying Chen

We propose an Exclusive Topic Modeling (ETM) for unsupervised text classification, which is able to 1) identify the field-specific keywords though less frequently appeared and 2) deliver well-structured topics with exclu…

text-classificationText ClassificationUnsupervised Text Classification

Concentrated Document Topic Model

2021-02-06 · Hao Lei, Ying Chen

We propose a Concentrated Document Topic Model(CDTM) for unsupervised text classification, which is able to produce a concentrated and sparse document topic distribution. In particular, an exponential entropy penalty is …

modeltext-classificationText ClassificationUnsupervised Text Classification

Learning Interpretable and Discrete Representations with Adversarial Training for Unsupervised Text Classification

2020-04-28 · Yau-Shian Wang, Hung-Yi Lee, Yun-Nung Chen

Learning continuous representations from unlabeled textual data has been increasingly studied for benefiting semi-supervised learning. Although it is relatively easier to interpret discrete representations, due to the di…

General Classificationtext-classificationText ClassificationUnsupervised Text Classification

Diversity-Based Generalization for Unsupervised Text Classification under Domain Shift

2020-02-25 · Jitin Krishnan, Hemant Purohit, Huzefa Rangwala

Domain adaptation approaches seek to learn from a source domain and generalize it to an unseen target domain. At present, the state-of-the-art unsupervised domain adaptation approaches for subjective text classification …

ClassificationDiversityDomain AdaptationGeneral Classification+4

Towards Unsupervised Text Classification Leveraging Experts and Word Embeddings

2019-07-01 · ACL 2019 7 · Zied Haj-Yahia, Adrien Sieg, L{\'e}a A. Deleris

Text classification aims at mapping documents into a set of predefined categories. Supervised machine learning models have shown great success in this area but they require a large number of labeled documents to reach ad…

ClassificationGeneral ClassificationText Categorizationtext-classification+3

Semantic Term "Blurring" and Stochastic "Barcoding" for Improved Unsupervised Text Classification

2018-11-06 · Robert Frank Martorano III

The abundance of text data being produced in the modern age makes it increasingly important to intuitively group, categorize, or classify text data by theme for efficient retrieval and search. Yet, the high dimensionalit…

ClusteringDocument ClassificationGeneral ClassificationRetrieval+4
1–14 / 14