TopicsRanksDC: Distance-based Topic Ranking applied on Two-Class Data
In this paper, we introduce a novel approach named TopicsRanksDC for topics ranking based on the distance between two clusters that are generated by each topic. We assume that our data consists of text documents that are associated with two-classes. Our approach ranks each topic contained in these text documents by its significance for separating the two-classes. Firstly, the algorithm detects topics using Latent Dirichlet Allocation (LDA). The words defining each topic are represented as two clusters, where each one is associated with one of the classes. We compute four distance metrics, Single Linkage, Complete Linkage, Average Linkage and distance between the centroid. We compare the results of LDA topics and random topics. The results show that the rank for LDA topics is much higher than random topics. The results of TopicsRanksDC tool are promising for future work to enable search engines to suggest related topics.
Code (0)
등록된 구현이 없습니다.
Tasks
Vocal Bursts Valence PredictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Re-Ranking Words to Improve Interpretability of Automatically Generated Topics
Topics models, such as LDA, are widely used in Natural Language Processing. Making their output interpretable is an important area of research with applications to areas such as the enhancement of exploratory search inte…
Interpretable Machine LearningRe-RankingRetrievalTime Series Analysis of Rankings: A GARCH-Type Approach
Ranking data are frequently obtained nowadays but there are still scarce methods for treating these data when temporally observed. The present paper contributes to this topic by proposing and developing novel models for …
Time SeriesTime Series AnalysisLinear Ranking Analysis
We extend the classical linear discriminant analysis (LDA) technique to linear ranking analysis (LRA), by considering the ranking order of classes centroids on the projected subspace. Under the constrain on the ranking o…
Zero-Shot LearningA Hypervolume Based Approach to Rank Intuitionistic Fuzzy Sets and Its Extension to Multi-criteria Decision Making Under Uncertainty
Ranking intuitionistic fuzzy sets with distance based ranking methods requires to calculate the distance between intuitionistic fuzzy set and a reference point which is known to have either maximum (positive ideal soluti…
Decision MakingDecision Making Under UncertaintyvalidGitRanking: A Ranking of GitHub Topics for Software Classification using Active Sampling
GitHub is the world's largest host of source code, with more than 150M repositories. However, most of these repositories are not labeled or inadequately so, making it harder for users to find relevant projects. There hav…
domain classificationSpecificity