paper-with-me

홈 › Papers

On Efficient Multilevel Clustering via Wasserstein Distances

2019-09-19 · Viet Huynh, Nhat Ho, Nhan Dam, XuanLong Nguyen, Mikhail Yurochkin, Hung Bui, and Dinh Phung

We propose a novel approach to the problem of multilevel clustering, which aims to simultaneously partition data in each group and discover grouping patterns among groups in a potentially large hierarchically structured corpus of data. Our method involves a joint optimization formulation over several spaces of discrete probability measures, which are endowed with Wasserstein distance metrics. We propose several variants of this problem, which admit fast optimization algorithms, by exploiting the connection to the problem of finding Wasserstein barycenters. Consistency properties are established for the estimates of both local and global clusters. Finally, experimental results with both synthetic and real data are presented to demonstrate the flexibility and scalability of the proposed approach.

📄 PDF Abstract BibTeX arXiv:1909.08787

Code (1)

viethhuynh/wasserstein-means 공식 구현

Tasks

Clustering

Similar Papers 제목 키워드 기반

Multilevel Clustering via Wasserstein Means

2017-06-13 · ICML 2017 8 · Nhat Ho, XuanLong Nguyen, Mikhail Yurochkin, Hung Hai Bui 외

We propose a novel approach to the problem of multilevel clustering, which aims to simultaneously partition data in each group and discover grouping patterns among groups in a potentially large hierarchically structured …

Clustering

Tree-Wasserstein Barycenter for Large-Scale Multilevel Clustering and Scalable Bayes

2019-10-10 · Tam Le, Viet Huynh, Nhat Ho, Dinh Phung 외

We study in this paper a variant of Wasserstein barycenter problem, which we refer to as tree-Wasserstein barycenter, by leveraging a specific class of ground metrics, namely tree metrics, for Wasserstein distance. Drawi…

Clustering

Fuzzy clustering of distribution-valued data using adaptive L2 Wasserstein distances

2016-05-02 · Antonio Irpino, Francisco De Carvalho, Rosanna Verde

Distributional (or distribution-valued) data are a new type of data arising from several sources and are considered as realizations of distributional variables. A new set of fuzzy c-means algorithms for data described by…

ClusteringVariable Selection

Wasserstein-based Kernels for Clustering: Application to Power Distribution Graphs

2025-03-18 · Alfredo Oneto, Blazhe Gjorgiev, Giovanni Sansavini

Many data clustering applications must handle objects that cannot be represented as vector data. In this context, the bag-of-vectors representation can be leveraged to describe complex objects through discrete distributi…

Clustering

Variational Wasserstein Clustering

2018-06-23 · ECCV 2018 9 · Liang Mi, Wen Zhang, Xianfeng GU, Yalin Wang

We propose a new clustering method based on optimal transportation. We solve optimal transportation with variational principles, and investigate the use of power diagrams as transportation plans for aggregating arbitrary…

ClusteringDomain AdaptationRepresentation Learning