paper-with-me

홈 › Papers

Hierarchical Clustering With Confidence

2025-12-06 · Di Wu, Jacob Bien, Snigdha Panigrahi arxiv

Agglomerative hierarchical clustering is one of the most widely used approaches for exploring how observations in a dataset relate to each other. However, its greedy nature makes it highly sensitive to small perturbations in the data, often producing different clustering results and making it difficult to separate genuine structure from spurious patterns. In this paper, we show how randomizing hierarchical clustering can be useful not just for measuring stability but also for designing valid hypothesis testing procedures based on the clustering results. We propose a simple randomization scheme together with a method for constructing a valid p-value at each node of the hierarchical clustering dendrogram that quantifies evidence against performing the greedy merge. Our test controls the Type I error rate, works with any hierarchical linkage without case-specific derivations, and simulations show it is substantially more powerful than existing selective inference approaches. To demonstrate the practical utility of our p-values, we develop an adaptive $α$-spending procedure that estimates the number of clusters, with a probabilistic guarantee on overestimation. Experiments on simulated and real data show that this estimate yields powerful clustering and can be used, for example, to assess clustering stability across multiple runs of the randomized algorithm.

📄 PDF Abstract BibTeX arXiv:2512.06522

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interactive Steering of Hierarchical Clustering

2020-09-21 · Weikai Yang, Xiting Wang, Jie Lu, Wenwen Dou 외

Hierarchical clustering is an important technique to organize big data for exploratory data analysis. However, existing one-size-fits-all hierarchical clustering methods often fail to meet the diverse needs of different …

Clustering

CoBaR: Confidence-Based Recommender

2018-08-21 · Fernando S. Aguiar Neto, Arthur F. da Costa, Marcelo G. Manzato

Neighborhood-based collaborative filtering algorithms usually adopt a fixed neighborhood size for every user or item, although groups of users or items may have different lengths depending on users' preferences. In this …

ClusteringCollaborative Filtering

Incorporating Annotator Uncertainty into Representations of Discourse Relations

2023-08-14 · S. Magalí López Cortez, Cassandra L. Jacobs

Annotation of discourse relations is a known difficult task, especially for non-expert annotators. In this paper, we investigate novice annotators' uncertainty on the annotation of discourse relations on spoken conversat…

ClusteringRelation

Hierarchical Graph Interaction Transformer with Dynamic Token Clustering for Camouflaged Object Detection

2024-08-27 · Siyuan Yao, Hao Sun, Tian-Zhu Xiang, Xiao Wang 외

Camouflaged object detection (COD) aims to identify the objects that seamlessly blend into the surrounding backgrounds. Due to the intrinsic similarity between the camouflaged objects and the background region, it is ext…

Decoderobject-detectionObject Detection

Towards Calibrated Deep Clustering Network

2024-03-04 · Yuheng Jia, Jianhong Cheng, Hui Liu, Junhui Hou

Deep clustering has exhibited remarkable performance; however, the over-confidence problem, i.e., the estimated confidence for a sample belonging to a particular cluster greatly exceeds its actual prediction accuracy, ha…

ClusteringDeep ClusteringPseudo Label