Pairwise proximity metrics for topic modelling evaluation based on BERT embeddings.
The use of topic modelling methods is a popular way to describe natural language text with a representative set of words. In order to evaluate such methods, objective metrics such as coherence and silhouette scores are commonly used. However, it has been shown that topic assessment based on such metrics does not align well with human judgment for classical document corpora such as articles, books and server logs and, at the same time, it is still unclear how appropriate they are for dialog data. In this paper, we investigate the most commonly used topic modelling evaluation scores in terms of their alignment with human judgment in the specific area of dialog speech. We show that there is still space for improvement in the objective evaluation of topic modelling, and propose a new group of metrics, called Pairwise Proximity metrics, that are shown to align better with human judgment, when compared to coherence and silhouette scores.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesSimilar Papers 제목 키워드 기반
Apples to Apples: A Systematic Evaluation of Topic Models
From statistical to neural models, a wide variety of topic modelling algorithms have been proposed in the literature. However, because of the diversity of datasets and metrics, there have not been many efforts to systema…
DiversityTopic ModelsWord EmbeddingsMetrics for describing dyadic movement: a review
In movement ecology, the few works that have taken collective behaviour into account are data-driven and rely on simplistic theoretical assumptions, relying in metrics that may or may not be measuring what is intended. I…
TopicNet: Making Additive Regularisation for Topic Modelling Accessible
This paper introduces TopicNet, a new Python module for topic modeling. This package, distributed under the MIT license, focuses on bringing additive regularization topic modelling (ARTM) to non-specialists using a gener…
Model SelectionOptimized Tracking of Topic Evolution
Topic evolution modeling has been researched for a long time and has gained considerable interest. A state-of-the-art method has been recently using word modeling algorithms in combination with community detection mechan…
Community DetectionA Data-driven Latent Semantic Analysis for Automatic Text Summarization using LDA Topic Modelling
With the advent and popularity of big data mining and huge text analysis in modern times, automated text summarization became prominent for extracting and retrieving important information from documents. This research in…
ArticlesExtractive SummarizationText Summarization