paper-with-me

홈 › Papers

Autoencoding Keyword Correlation Graph for Document Clustering

2020-07-01 · ACL 2020 6 · Billy Chiu, Sunil Kumar Sahu, Derek Thomas, Neha Sengupta, Mohammady Mahdy

Document clustering requires a deep understanding of the complex structure of long-text; in particular, the intra-sentential (local) and inter-sentential features (global). Existing representation learning models do not fully capture these features. To address this, we present a novel graph-based representation for document clustering that builds a \textit{graph autoencoder} (GAE) on a Keyword Correlation Graph. The graph is constructed with topical keywords as nodes and multiple local and global features as edges. A GAE is employed to aggregate the two sets of features by learning a latent representation which can jointly reconstruct them. Clustering is then performed on the learned representations, using vector dimensions as features for inducing document classes. Extensive experiments on two datasets show that the features learned by our approach can achieve better clustering performance than other existing features, including term frequency-inverse document frequency and average embedding.

📄 PDF Abstract BibTeX

Code (1)

1997alireza/Autoencoding-Graph-for-Document-Clustering

Tasks

ClusteringRepresentation Learning

Similar Papers 제목 키워드 기반

Semantic Document Clustering on Named Entity Features

2018-07-20 · Cao Tru H., Ngo Vuong M., Hong Dung T., Quan Tho T.

Keyword-based information processing has limitations due to simple treatment of words. In this paper, we introduce named entities as objectives into document clustering, which are the key elements defining document seman…

Clustering

Leveraging web resources for keyword assignment to short text documents

2017-06-19 · Singhal Ayush, Kasturi Ravindra, Sharma Ankit, Srivastava Jaideep

Assigning relevant keywords to documents is very important for efficient retrieval, clustering and management of the documents. Especially with the web corpus deluged with digital documents, automation of this task is of…

Keyword ExtractionManagementRetrieval

Higher-Order Correlation Clustering for Image Segmentation

2011-12-01 · NeurIPS 2011 12 · Sungwoong Kim, Sebastian Nowozin, Pushmeet Kohli, Chang D. Yoo

For many of the state-of-the-art computer vision algorithms, image segmentation is an important preprocessing step. As such, several image segmentation algorithms have been proposed, however, with certain reservation due…

Clusteringgraph partitioningImage SegmentationSegmentation+2

A Bibliographic View on Constrained Clustering

2022-09-22 · Ludmila Kuncheva, Francis Williams, Samuel Hennessey

A keyword search on constrained clustering on Web-of-Science returned just under 3,000 documents. We ran automatic analyses of those, and compiled our own bibliography of 183 papers which we analysed in more detail based…

Active LearningClusteringConstrained ClusteringEnsemble Learning

Document Network Projection in Pretrained Word Embedding Space

2020-01-16 · Antoine Gourru, Adrien Guille, Julien Velcin, Julien Jacques

We present Regularized Linear Embedding (RLE), a novel method that projects a collection of linked documents (e.g. citation network) into a pretrained word embedding space. In addition to the textual content, we leverage…

ClusteringGeneral ClassificationInformation RetrievalLink Prediction+3