paper-with-me

홈 › Papers

ERICA: Quantifying Replicability of Cluster Analysis

2026-05-29 · Siamak K. Sorooshyari, Manuel A. Rivas, Robert Tibshirani arxiv

Despite being ubiquitous in science, clustering lacks a unified framework for quantitatively evaluating the replicability of its results. We present evaluating replicability via iterative clustering assignments (ERICA), a method for determining whether clusters can be identified reproducibly in a dataset. The pipeline computes a statistic that determines whether reproducible cluster structure is present in a dataset. Quantitative visualization methods are also introduced to characterize similarities between clusters and identify observations that may represent outliers or unstable assignments. Experiments on synthetic datasets demonstrate that ERICA successfully identifies reproducible cluster structure. In contrast, application of ERICA to three breast cancer gene-expression datasets reveals instances in which clustering solutions are not reproducible. The study underscores the importance of rigorously evaluating clustering solutions and provides a practical framework for doing so.

📄 PDF Abstract BibTeX arXiv:2606.00302

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GeoAI Reproducibility and Replicability: a computational and spatial perspective

2024-04-15 · Wenwen Li, Chia-Yu Hsu, Sizhe Wang, Peter Kedron

GeoAI has emerged as an exciting interdisciplinary research area that combines spatial theories and data with cutting-edge AI models to address geospatial problems in a novel, data-driven manner. While GeoAI research has…

An Initial Seed Selection Algorithm for K-means Clustering of Georeferenced Data to Improve Replicability of Cluster Assignments for Mapping Application

2016-04-17 · Fouad Khan

K-means is one of the most widely used clustering algorithms in various disciplines, especially for large datasets. However the method is known to be highly sensitive to initial seed selection of cluster centers. K-means…

AttributeClusteringComputational Efficiency

The Cost of Replicability in Active Learning

2024-12-12 · Rupkatha Hira, Dominik Kau, Jessica Sorrell

Active learning aims to reduce the required number of labeled data for machine learning algorithms by selectively querying the labels of initially unlabeled data points. Ensuring the replicability of results, where an al…

Active Learning

Sensitivity of Stability: Theoretical & Empirical Analysis of Replicability for Adaptive Data Selection in Transfer Learning

2025-08-06 · Prabhav Singh, Jessica Sorrell arxiv

The widespread adoption of transfer learning has revolutionized machine learning by enabling efficient adaptation of pre-trained models to new domains. However, the reliability of these adaptations remains poorly underst…

Transfer Learning

From Model Performance to Claim: How a Change of Focus in Machine Learning Replicability Can Help Bridge the Responsibility Gap

2024-04-19 · Tianqi Kou

Two goals - improving replicability and accountability of Machine Learning research respectively, have accrued much attention from the AI ethics and the Machine Learning community. Despite sharing the measures of improvi…

Ethics