paper-with-me

홈 › Papers

A partition-based similarity for classification distributions

2020-11-12 · Hayden S. Helm, Ronak D. Mehta, Brandon Duderstadt, Weiwei Yang, Christoper M. White, Ali Geisa, Joshua T. Vogelstein, Carey E. Priebe

Herein we define a measure of similarity between classification distributions that is both principled from the perspective of statistical pattern recognition and useful from the perspective of machine learning practitioners. In particular, we propose a novel similarity on classification distributions, dubbed task similarity, that quantifies how an optimally-transformed optimal representation for a source distribution performs when applied to inference related to a target distribution. The definition of task similarity allows for natural definitions of adversarial and orthogonal distributions. We highlight limiting properties of representations induced by (universally) consistent decision rules and demonstrate in simulation that an empirical estimate of task similarity is a function of the decision rule deployed for inference. We demonstrate that for a given target distribution, both transfer efficiency and semantic similarity of candidate source distributions correlate with empirical task similarity.

📄 PDF Abstract BibTeX arXiv:2011.06557

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

Discriminative Similarity for Data Clustering

2021-09-17 · ICLR 2022 4 · Yingzhen Yang, Ping Li

Similarity-based clustering methods separate data into clusters according to the pairwise similarity between the data, and the pairwise similarity is crucial for their performance. In this paper, we propose {\em Clusteri…

Clustering

Online hierarchical partitioning of the output space in extreme multi-label data stream

2025-07-28 · Lara Neves, Afonso Lourenço, Alberto Cano, Goreti Marreiros arxiv

Mining data streams with multi-label outputs poses significant challenges due to evolving distributions, high-dimensional label spaces, sparse label occurrences, and complex label dependencies. Moreover, concept drift af…

Multi-Label ClassificationMulti-Label Learning

On a Theory of Nonparametric Pairwise Similarity for Clustering: Connecting Clustering to Classification

2014-12-01 · NeurIPS 2014 12 · Yingzhen Yang, Feng Liang, Shuicheng Yan, Zhangyang Wang 외

Pairwise clustering methods partition the data space into clusters by the pairwise similarity between data points. The success of pairwise clustering largely depends on the pairwise similarity function defined over the d…

ClusteringDensity EstimationGeneral ClassificationMulti-class Classification

Nonparametric Hierarchical Clustering of Functional Data

2014-07-02 · Marc Boullé, Romain Guigourès, Fabrice Rossi

In this paper, we deal with the problem of curves clustering. We propose a nonparametric method which partitions the curves into clusters and discretizes the dimensions of the curve points into intervals. The cross-produ…

ClusteringModel Selection

Wasserstein $K$-means for clustering probability distributions

2022-09-14 · Yubo Zhuang, Xiaohui Chen, Yun Yang

Clustering is an important exploratory data analysis technique to group objects based on their similarity. The widely used $K$-means clustering method relies on some notion of distance to partition data into a fewer numb…

Clustering