Papers Short Text Clustering
“Short Text Clustering” 태그가 달린 논문 39편 · 필터 해제
Beyond Looking Up, Try Looking Around: Harmonizing Global Structure and Local Consistency in Optimal Transport for Short Text Clustering
Pseudo-labeling based on Optimal Transport (OT) has become an effective mechanism for enhancing short text clustering. Existing OT methods are short in modeling semantic consistencies between samples, which may assign di…
Short Text ClusteringLUMI: Unsupervised Intent Clustering with Multiple Pseudo-Labels
In this paper, we propose an intuitive, training-free and label-free method for intent clustering in conversational search. Current approaches to short text clustering use LLM-generated pseudo-labels to enrich text repre…
Short Text ClusteringLLMs Enable Bag-of-Texts Representations for Short-Text Clustering
In this paper, we propose a training-free method for unsupervised short text clustering that relies less on careful selection of embedders than other methods. In customer-facing chatbots, companies are dealing with large…
Short Text ClusteringNormalisation of SWIFT Message Counterparties with Feature Extraction and Clustering
Short text clustering is a known use case in the text analytics community. When the structure and content falls in the natural language domain e.g. Twitter posts or instant messages, then natural language techniques can …
Short Text ClusteringAn Enhanced Model-based Approach for Short Text Clustering
Short text clustering has become increasingly important with the popularity of social media like Twitter, Google+, and Facebook. Existing methods can be broadly categorized into two paradigms: topic model-based approache…
Representation LearningShort Text ClusteringMoving Past Single Metrics: Exploring Short-Text Clustering Across Multiple Resolutions
Cluster number is typically a parameter selected at the outset in clustering problems, and while impactful, the choice can often be difficult to justify. Inspired by bioinformatics, this study examines how the nature of …
ClusteringInformativenessShort Text ClusteringText ClusteringReliable Pseudo-labeling via Optimal Transport with Attention for Short Text Clustering
Short text clustering has gained significant attention in the data mining community. However, the limited valuable information contained in short texts often leads to low-discriminative representations, increasing the di…
ClusteringContrastive LearningRepresentation LearningShort Text Clustering+1Discriminative Representation learning via Attention-Enhanced Contrastive Learning for Short Text Clustering
Contrastive learning has gained significant attention in short text clustering, yet it has an inherent drawback of mistakenly identifying samples from the same category as negatives and then separating them in the featur…
ClusteringContrastive LearningPseudo LabelRepresentation Learning+2Hierarchical mixtures of Unigram models for short text clustering: The role of Beta-Liouville priors
This paper presents a variant of the Multinomial mixture model tailored to the unsupervised classification of short text data. While the Multinomial probability vector is traditionally assigned a Dirichlet prior distribu…
Short Text ClusteringText ClusteringVariational InferenceExtracting Sentence Embeddings from Pretrained Transformer Models
Pre-trained transformer models shine in many natural language processing tasks and therefore are expected to bear the representation of the input sentence or text meaning. These sentence-level embeddings are also importa…
ClusteringRetrieval-augmented GenerationSemantic Textual SimilaritySentence+6Guiding Sentiment Analysis with Hierarchical Text Clustering: Analyzing the German X/Twitter Discourse on Face Masks in the 2020 COVID-19 Pandemic
Social media are a critical component of the information ecosystem during public health crises. Understanding the public discourse is essential for effective communication and misinformation mitigation. Computational met…
ClusteringData VisualizationHierarchical Text ClusteringMisinformation+4Human-interpretable clustering of short-text using large language models
Clustering short text is a difficult problem, due to the low word co-occurrence between short text documents. This work shows that large language models (LLMs) can overcome the limitations of traditional clustering appro…
ClusteringShort Text ClusteringText ClusteringFederated Learning for Short Text Clustering
Short text clustering has been popularly studied for its significance in mining valuable insights from many short texts. In this paper, we focus on the federated short text clustering (FSTC) problem, i.e., clustering sho…
ClusteringFederated LearningShort Text ClusteringText ClusteringRobust Representation Learning with Reliable Pseudo-labels Generation via Self-Adaptive Optimal Transport for Short Text Clustering
Short text clustering is challenging since it takes imbalanced and noisy data as inputs. Existing approaches cannot solve this problem well, since (1) they are prone to obtain degenerate solutions especially on heavy imb…
ClusteringContrastive LearningPseudo LabelRepresentation Learning+2CEIL: A General Classification-Enhanced Iterative Learning Framework for Text Clustering
Text clustering, as one of the most fundamental challenges in unsupervised learning, aims at grouping semantically similar text segments without relying on human annotations. With the rapid development of deep learning, …
ClusteringDeep ClusteringGeneral ClassificationLanguage Modeling+4Twin Contrastive Learning for Online Clustering
This paper proposes to perform online clustering by conducting twin contrastive learning (TCL) at the instance and cluster level. Specifically, we find that when the data is projected into a feature space with a dimensio…
ClusteringContrastive LearningDeep ClusteringImage Clustering+2Improving Deep Embedded Clustering via Learning Cluster-level Representations
Driven by recent advances in neural networks, various Deep Embedding Clustering (DEC) based short text clustering models are being developed. In these works, latent representation learning and text clustering are perform…
ClusteringContrastive LearningRepresentation LearningShort Text Clustering+1EASE: Entity-Aware Contrastive Learning of Sentence Embedding
We present EASE, a novel method for learning sentence embeddings via contrastive learning between sentences and their related entities. The advantage of using entity supervision is twofold: (1) entities have been shown t…
ClusteringContrastive LearningSemantic Textual SimilaritySentence+6EASE: Entity-Aware Contrastive Learning of Sentence Embedding
We present EASE, a novel method for learning sentence embeddings via contrastive learning between sentences and their related entities.The advantage of using entity supervision is twofold: (1) entities have been shown to…
ClusteringContrastive LearningSemantic Textual SimilaritySentence+6Representation Learning for Short Text Clustering
Effective representation learning is critical for short text clustering due to the sparse, high-dimensional and noise attributes of short text corpus. Existing pre-trained models (e.g., Word2vec and BERT) have greatly im…
ClusteringRepresentation LearningShort Text ClusteringText Clustering