paper-with-me

홈 › Papers

Influence of various text embeddings on clustering performance in NLP

2023-05-04 · Rohan Saha

With the advent of e-commerce platforms, reviews are crucial for customers to assess the credibility of a product. The star ratings do not always match the review text written by the customer. For example, a three star rating (out of five) may be incongruous with the review text, which may be more suitable for a five star review. A clustering approach can be used to relabel the correct star ratings by grouping the text reviews into individual groups. In this work, we explore the task of choosing different text embeddings to represent these reviews and also explore the impact the embedding choice has on the performance of various classes of clustering algorithms. We use contextual (BERT) and non-contextual (Word2Vec) text embeddings to represent the text and measure their impact of three classes on clustering algorithms - partitioning based (KMeans), single linkage agglomerative hierarchical, and density based (DBSCAN and HDBSCAN), each with various experimental settings. We use the silhouette score, adjusted rand index score, and cluster purity score metrics to evaluate the performance of the algorithms and discuss the impact of different embeddings on the clustering performance. Our results indicate that the type of embedding chosen drastically affects the performance of the algorithm, the performance varies greatly across different types of clustering algorithms, no embedding type is better than the other, and DBSCAN outperforms KMeans and single linkage agglomerative clustering but also labels more data points as outliers. We provide a thorough comparison of the performances of different algorithms and provide numerous ideas to foster further research in the domain of text clustering.

📄 PDF Abstract BibTeX arXiv:2305.03144

Code (1)

simpleparadox/cmput_697_project 공식 구현 pytorch

Tasks

ClusteringText Clustering

Similar Papers 제목 키워드 기반

Text Clustering with Large Language Model Embeddings

2024-03-22 · Alina Petukhova, João P. Matos-Carvalho, Nuno Fachada

Text clustering is an important method for organising the increasing volume of digital content, aiding in the structuring and discovery of hidden patterns in uncategorised data. The effectiveness of text clustering large…

ClusteringDimensionality ReductionLanguage ModelingLanguage Modelling+3

Word Embeddings and Validity Indexes in Fuzzy Clustering

2022-04-26 · Danial Toufani-Movaghar, Mohammad-Reza Feizi-Derakhshi

In the new era of internet systems and applications, a concept of detecting distinguished topics from huge amounts of text has gained a lot of attention. These methods use representation of text in a numerical format -- …

ClusteringSemantic SimilaritySemantic Textual SimilarityWord Embeddings

Measure and Evaluation of Semantic Divergence across Two Languages

2021-08-01 · ACL 2021 5 · Syrielle Montariol, Alexandre Allauzen

Languages are dynamic systems: word usage may change over time, reflecting various societal factors. However, all languages do not evolve identically: the impact of an event, the influence of a trend or thinking, can dif…

TranslationVocal Bursts Valence PredictionWord Embeddings

Homonym Identification using BERT -- Using a Clustering Approach

2021-01-07 · Rohan Saha

Homonym identification is important for WSD that require coarse-grained partitions of senses. The goal of this project is to determine whether contextual information is sufficient for identifying a homonymous word. To ca…

Clustering

Error Discovery by Clustering Influence Embeddings

2023-12-07 · NeurIPS 2023 11 · Fulton Wang, Julius Adebayo, Sarah Tan, Diego Garcia-Olano 외

We present a method for identifying groups of test examples -- slices -- on which a model under-performs, a task now known as slice discovery. We formalize coherence -- a requirement that erroneous predictions, within a …

ClusteringSlice Discovery