paper-with-me

Papers Short Text Clustering

“Short Text Clustering” 태그가 달린 논문 39편 · 필터 해제

Beyond Looking Up, Try Looking Around: Harmonizing Global Structure and Local Consistency in Optimal Transport for Short Text Clustering

2026-07-12 · Zhihao Yao, Yuxuan Gu, Jixuan Yin, Bo Li arxiv

Pseudo-labeling based on Optimal Transport (OT) has become an effective mechanism for enhancing short text clustering. Existing OT methods are short in modeling semantic consistencies between samples, which may assign di…

Short Text Clustering

LUMI: Unsupervised Intent Clustering with Multiple Pseudo-Labels

2025-10-16 · I-Fan Lin, Faegheh Hasibi, Suzan Verberne arxiv

In this paper, we propose an intuitive, training-free and label-free method for intent clustering in conversational search. Current approaches to short text clustering use LLM-generated pseudo-labels to enrich text repre…

Short Text Clustering

LLMs Enable Bag-of-Texts Representations for Short-Text Clustering

2025-10-08 · I-Fan Lin, Faegheh Hasibi, Suzan Verberne arxiv

In this paper, we propose a training-free method for unsupervised short text clustering that relies less on careful selection of embedders than other methods. In customer-facing chatbots, companies are dealing with large…

Short Text Clustering

Normalisation of SWIFT Message Counterparties with Feature Extraction and Clustering

2025-08-24 · Thanasis Schoinas, Benjamin Guinard, Diba Esbati, Richard Chalk arxiv

Short text clustering is a known use case in the text analytics community. When the structure and content falls in the natural language domain e.g. Twitter posts or instant messages, then natural language techniques can …

Short Text Clustering

An Enhanced Model-based Approach for Short Text Clustering

2025-07-18 · Enhao Cheng, Shoujia Zhang, Jianhua Yin, Xuemeng Song 외 arxiv

Short text clustering has become increasingly important with the popularity of social media like Twitter, Google+, and Facebook. Existing methods can be broadly categorized into two paradigms: topic model-based approache…

Representation LearningShort Text Clustering

Moving Past Single Metrics: Exploring Short-Text Clustering Across Multiple Resolutions

2025-02-24 · Justin Miller, Tristram Alexander

Cluster number is typically a parameter selected at the outset in clustering problems, and while impactful, the choice can often be difficult to justify. Inspired by bioinformatics, this study examines how the nature of …

ClusteringInformativenessShort Text ClusteringText Clustering

Reliable Pseudo-labeling via Optimal Transport with Attention for Short Text Clustering

2025-01-25 · Zhihao Yao, Jixuan Yin, Bo Li

Short text clustering has gained significant attention in the data mining community. However, the limited valuable information contained in short texts often leads to low-discriminative representations, increasing the di…

ClusteringContrastive LearningRepresentation LearningShort Text Clustering+1

Discriminative Representation learning via Attention-Enhanced Contrastive Learning for Short Text Clustering

2025-01-07 · Zhihao Yao

Contrastive learning has gained significant attention in short text clustering, yet it has an inherent drawback of mistakenly identifying samples from the same category as negatives and then separating them in the featur…

ClusteringContrastive LearningPseudo LabelRepresentation Learning+2

Hierarchical mixtures of Unigram models for short text clustering: The role of Beta-Liouville priors

2024-10-29 · Massimo Bilancia, Samuele Magro

This paper presents a variant of the Multinomial mixture model tailored to the unsupervised classification of short text data. While the Multinomial probability vector is traditionally assigned a Dirichlet prior distribu…

Short Text ClusteringText ClusteringVariational Inference

Extracting Sentence Embeddings from Pretrained Transformer Models

2024-08-15 · Lukas Stankevičius, Mantas Lukoševičius

Pre-trained transformer models shine in many natural language processing tasks and therefore are expected to bear the representation of the input sentence or text meaning. These sentence-level embeddings are also importa…

ClusteringRetrieval-augmented GenerationSemantic Textual SimilaritySentence+6

Guiding Sentiment Analysis with Hierarchical Text Clustering: Analyzing the German X/Twitter Discourse on Face Masks in the 2020 COVID-19 Pandemic

2024-08-01 · Proceedings of the 14th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis 2024 8 · Silvan Wehrli, Chisom Ezekannagha, Georges Hattab, Tamara Boender 외

Social media are a critical component of the information ecosystem during public health crises. Understanding the public discourse is essential for effective communication and misinformation mitigation. Computational met…

ClusteringData VisualizationHierarchical Text ClusteringMisinformation+4

Human-interpretable clustering of short-text using large language models

2024-05-12 · Justin K. Miller, Tristram J. Alexander

Clustering short text is a difficult problem, due to the low word co-occurrence between short text documents. This work shows that large language models (LLMs) can overcome the limitations of traditional clustering appro…

ClusteringShort Text ClusteringText Clustering

Federated Learning for Short Text Clustering

2023-11-23 · Mengling Hu, Chaochao Chen, Weiming Liu, Xinting Liao 외

Short text clustering has been popularly studied for its significance in mining valuable insights from many short texts. In this paper, we focus on the federated short text clustering (FSTC) problem, i.e., clustering sho…

ClusteringFederated LearningShort Text ClusteringText Clustering

Robust Representation Learning with Reliable Pseudo-labels Generation via Self-Adaptive Optimal Transport for Short Text Clustering

2023-05-23 · Xiaolin Zheng, Mengling Hu, Weiming Liu, Chaochao Chen 외

Short text clustering is challenging since it takes imbalanced and noisy data as inputs. Existing approaches cannot solve this problem well, since (1) they are prone to obtain degenerate solutions especially on heavy imb…

ClusteringContrastive LearningPseudo LabelRepresentation Learning+2

CEIL: A General Classification-Enhanced Iterative Learning Framework for Text Clustering

2023-04-20 · Mingjun Zhao, Mengzhen Wang, Yinglong Ma, Di Niu 외

Text clustering, as one of the most fundamental challenges in unsupervised learning, aims at grouping semantically similar text segments without relying on human annotations. With the rapid development of deep learning, …

ClusteringDeep ClusteringGeneral ClassificationLanguage Modeling+4

Twin Contrastive Learning for Online Clustering

2022-10-21 · Yunfan Li, Mouxing Yang, Dezhong Peng, Taihao Li 외

This paper proposes to perform online clustering by conducting twin contrastive learning (TCL) at the instance and cluster level. Specifically, we find that when the data is projected into a feature space with a dimensio…

ClusteringContrastive LearningDeep ClusteringImage Clustering+2

Improving Deep Embedded Clustering via Learning Cluster-level Representations

2022-10-01 · COLING 2022 10 · Qing Yin, Zhihua Wang, Yunya Song, Yida Xu 외

Driven by recent advances in neural networks, various Deep Embedding Clustering (DEC) based short text clustering models are being developed. In these works, latent representation learning and text clustering are perform…

ClusteringContrastive LearningRepresentation LearningShort Text Clustering+1

EASE: Entity-Aware Contrastive Learning of Sentence Embedding

2022-05-09 · NAACL 2022 7 · Sosuke Nishikawa, Ryokan Ri, Ikuya Yamada, Yoshimasa Tsuruoka 외

We present EASE, a novel method for learning sentence embeddings via contrastive learning between sentences and their related entities. The advantage of using entity supervision is twofold: (1) entities have been shown t…

ClusteringContrastive LearningSemantic Textual SimilaritySentence+6

EASE: Entity-Aware Contrastive Learning of Sentence Embedding

2022-01-16 · ACL ARR January 2022 1 · Anonymous

We present EASE, a novel method for learning sentence embeddings via contrastive learning between sentences and their related entities.The advantage of using entity supervision is twofold: (1) entities have been shown to…

ClusteringContrastive LearningSemantic Textual SimilaritySentence+6

Representation Learning for Short Text Clustering

2021-09-21 · Hui Yin, XiangYu Song, Shuiqiao Yang, Guangyan Huang 외

Effective representation learning is critical for short text clustering due to the sparse, high-dimensional and noise attributes of short text corpus. Existing pre-trained models (e.g., Word2vec and BERT) have greatly im…

ClusteringRepresentation LearningShort Text ClusteringText Clustering
1–20 / 39 다음 →