Leveraging Meta Information in Short Text Aggregation
Short texts such as tweets often contain insufficient word co-occurrence information for training conventional topic models. To deal with the insufficiency, we propose a generative model that aggregates short texts into clusters by leveraging the associated meta information. Our model can generate more interpretable topics as well as document clusters. We develop an effective Gibbs sampling algorithm favoured by the fully local conjugacy in the model. Extensive experiments demonstrate that our model achieves better performance in terms of document clustering and topic coherence.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringTopic ModelsSimilar Papers 제목 키워드 기반
COEM: Cross-Modal Embedding for MetaCell Identification
Metacells are disjoint and homogeneous groups of single-cell profiles, representing discrete and highly granular cell states. Existing metacell algorithms tend to use only one modality to infer metacells, even though sin…
Meta-Path-based Fake News Detection Leveraging Multi-level Social Context Information
Fake news, false or misleading information presented as news, has a significant impact on many aspects of society, such as in politics or healthcare domains. Due to the deceiving nature of fake news, applying Natural Lan…
Fake News DetectionStance DetectionShort-Term and Long-Term Context Aggregation Network for Video Inpainting
Video inpainting aims to restore missing regions of a video and has many applications such as video editing and object removal. However, existing methods either suffer from inaccurate short-term context aggregation or ra…
Video EditingVideo InpaintingMoviescope: Large-scale Analysis of Movies using Multiple Modalities
Film media is a rich form of artistic expression. Unlike photography, and short videos, movies contain a storyline that is deliberately complex and intricate in order to engage its audience. In this paper we present a la…
GloCOM: A Short Text Neural Topic Model via Global Clustering Context
Uncovering hidden topics from short texts is challenging for traditional and neural models due to data sparsity, which limits word co-occurrence patterns, and label sparsity, stemming from incomplete reconstruction targe…
ClusteringTopic Models