Elastic deep autoencoder for text embedding clustering by an improved graph regularization
Text clustering is a task for grouping extracted information of the text in different clusters, which has many applications in recommender systems, sentiment analysis, and more. Deep learning-based methods have become increasingly popular due to their high accuracy in identifying nonlinear structures. They usually consist of two major parts: dimensionality reduction and clustering. Autoencoders are simple unsupervised neural networks used for better representation of low-dimensional data and have shown good performance in dealing with non-linear features. However, while they utilize the Frobenius norm to deal well with Gaussian noise, they are sensitive to outlier data and Laplacian noise. In this paper, a deep autoencoder with an adapted elastic loss for text embedding clustering (EDA-TEC) is proposed. The elastic loss is a combination of the Frobenius norm and L2,1-norm to consider both types of noises. Additionally, to maintain the high-dimensional data geometric structure, a modified graph regularization term based on the weighted cosine similarity measure is used. EDA-TEC also improves clustering results by considering the sparsity regularization of the manifold representation data. In this jointly end-to-end deep learning model, better representation and text clustering results are achieved with high accuracy on common datasets compared to existing methods.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringDimensionality ReductionRecommendation SystemsSentiment AnalysisText ClusteringSimilar Papers 제목 키워드 기반
Masked AutoEncoder for Graph Clustering without Pre-defined Cluster Number k
Graph clustering algorithms with autoencoder structures have recently gained popularity due to their efficient performance and low training cost. However, for existing graph autoencoder clustering algorithms based on GCN…
ClusteringDecoderGraph ClusteringSDEC: Semantic Deep Embedded Clustering
The high dimensional and semantically complex nature of textual Big data presents significant challenges for text clustering, which frequently lead to suboptimal groupings when using conventional techniques like k-means …
Text ClusteringDeep Learning-Based Approach for Improving Relational Aggregated Search
Due to an information explosion on the internet, there is a need for the development of aggregated search systems that can boost the retrieval and management of content in various formats. To further improve the clusteri…
Representation LearningSparse Elasticity Reconstruction and Clustering using Local Displacement Fields
This paper introduces an elasticity reconstruction method based on local displacement observations of elastic bodies. Sparse reconstruction theory is applied to formulate the underdetermined inverse problems of elasticit…
ClusteringClustering Time Series Data with Gaussian Mixture Embeddings in a Graph Autoencoder Framework
Time series data analysis is prevalent across various domains, including finance, healthcare, and environmental monitoring. Traditional time series clustering methods often struggle to capture the complex temporal depend…
ClusteringManagementPortfolio OptimizationTime Series+1