Learning Taxonomies by Dependence Maximization
We introduce a family of unsupervised algorithms, numerical taxonomy clustering, to simultaneously cluster data, and to learn a taxonomy that encodes the relationship between the clusters. The algorithms work by maximizing the dependence between the taxonomy and the original data. The resulting taxonomy is a more informative visualization of complex data than simple clustering; in addition, taking into account the relations between different clusters is shown to substantially improve the quality of the clustering, when compared with state-of-the-art algorithms in the literature (both spectral clustering and a previous dependence maximization approach). We demonstrate our algorithm on image and text data.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
On the Limitation of Kernel Dependence Maximization for Feature Selection
A simple and intuitive method for feature selection consists of choosing the feature subset that maximizes a nonparametric measure of dependence between the response and the features. A popular proposal from the literatu…
feature selectionExpected Utility Maximization and Conditional Value-at-Risk Deviation-based Sharpe Ratio in Dynamic Stochastic Portfolio Optimization
In this paper we investigate the expected terminal utility maximization approach for a dynamic stochastic portfolio optimization problem. We solve it numerically by solving an evolutionary Hamilton-Jacobi-Bellman equatio…
Portfolio OptimizationUsing Zero-shot Prompting in the Automatic Creation and Expansion of Topic Taxonomies for Tagging Retail Banking Transactions
This work presents an unsupervised method for automatically constructing and expanding topic taxonomies using instruction-based fine-tuned LLMs (Large Language Models). We apply topic modeling and keyword extraction tech…
Keyword ExtractionHierarchical Text Classification with LLM-Refined Taxonomies
Hierarchical text classification (HTC) depends on taxonomies that organize labels into structured hierarchies. However, many real-world taxonomies introduce ambiguities, such as identical leaf names under similar parent …
Text ClassificationNo Gold Standard, No Problem: Reference-Free Evaluation of Taxonomies
We introduce two reference-free metrics for quality evaluation of taxonomies. The first metric evaluates robustness by calculating the correlation between semantic and taxonomic similarity, covering a type of error not h…
Natural Language Inference