Autoencoding Keyword Correlation Graph for Document Clustering
Document clustering requires a deep understanding of the complex structure of long-text; in particular, the intra-sentential (local) and inter-sentential features (global). Existing representation learning models do not fully capture these features. To address this, we present a novel graph-based representation for document clustering that builds a \textit{graph autoencoder} (GAE) on a Keyword Correlation Graph. The graph is constructed with topical keywords as nodes and multiple local and global features as edges. A GAE is employed to aggregate the two sets of features by learning a latent representation which can jointly reconstruct them. Clustering is then performed on the learned representations, using vector dimensions as features for inducing document classes. Extensive experiments on two datasets show that the features learned by our approach can achieve better clustering performance than other existing features, including term frequency-inverse document frequency and average embedding.
Code (1)
Tasks
ClusteringRepresentation LearningSimilar Papers 제목 키워드 기반
Semantic Document Clustering on Named Entity Features
Keyword-based information processing has limitations due to simple treatment of words. In this paper, we introduce named entities as objectives into document clustering, which are the key elements defining document seman…
ClusteringLeveraging web resources for keyword assignment to short text documents
Assigning relevant keywords to documents is very important for efficient retrieval, clustering and management of the documents. Especially with the web corpus deluged with digital documents, automation of this task is of…
Keyword ExtractionManagementRetrievalHigher-Order Correlation Clustering for Image Segmentation
For many of the state-of-the-art computer vision algorithms, image segmentation is an important preprocessing step. As such, several image segmentation algorithms have been proposed, however, with certain reservation due…
Clusteringgraph partitioningImage SegmentationSegmentation+2A Bibliographic View on Constrained Clustering
A keyword search on constrained clustering on Web-of-Science returned just under 3,000 documents. We ran automatic analyses of those, and compiled our own bibliography of 183 papers which we analysed in more detail based…
Active LearningClusteringConstrained ClusteringEnsemble LearningDocument Network Projection in Pretrained Word Embedding Space
We present Regularized Linear Embedding (RLE), a novel method that projects a collection of linked documents (e.g. citation network) into a pretrained word embedding space. In addition to the textual content, we leverage…
ClusteringGeneral ClassificationInformation RetrievalLink Prediction+3