Can Topic Modelling benefit from Word Sense Information?
This paper proposes a new topic model that exploits word sense information in order to discover less redundant and more informative topics. Word sense information is obtained from WordNet and the discovered topics are groups of synsets, instead of mere surface words. A key feature is that all the known senses of a word are considered, with their probabilities. Alternative configurations of the model are described and compared to each other and to LDA, the most popular topic model. However, the obtained results suggest that there are no benefits of enriching LDA with word sense information.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
unimelb: Topic Modelling-based Word Sense Induction
unimelb: Topic Modelling-based Word Sense Induction for Web Snippet Clustering
Multivariate Gaussian Topic Modelling: A novel approach to discover topics with greater semantic coherence
An important aspect of text mining involves information retrieval in form of discovery of semantic themes (topics) from documents using topic modelling. While generative topic models like Latent Dirichlet Allocation (LDA…
Information RetrievalTopic ModelsMetaLDA: a Topic Model that Efficiently Incorporates Meta information
Besides the text content, documents and their associated words usually come with rich sets of meta informa- tion, such as categories of documents and semantic/syntactic features of words, like those encoded in word embed…
Topic ModelsWord EmbeddingsWord Sense Induction with Knowledge Distillation from BERT
Pre-trained contextual language models are ubiquitously employed for language understanding tasks, but are unsuitable for resource-constrained systems. Noncontextual word embeddings are an efficient alternative in these …
Knowledge DistillationLanguage ModelingLanguage ModellingWord Embeddings+2