paper-with-me

Papers

Two-cluster test

2025-07-11 · Xinying Liu, Lianyu Hu, Mudi Jiang, Simeng Zhang, Jun Lou, Zengyou He arxiv

Cluster analysis is a fundamental research issue in statistics and machine learning. In many modern clustering methods, we need to determine whether two subsets of samples come from the same cluster. Since these subsets are usually generated by certain clustering procedures, the deployment of classic two-sample tests in this context would yield extremely smaller p-values, leading to inflated Type-I error rate. To overcome this bias, we formally introduce the two-cluster test issue and argue that it is a totally different significance testing issue from conventional two-sample test. Meanwhile, we present a new method based on the boundary points between two subsets to derive an analytical p-value for the purpose of significance quantification. Experiments on both synthetic and real data sets show that the proposed test is able to significantly reduce the Type-I error rate, in comparison with several classic two-sample testing methods. More importantly, the practical usage of such two-cluster test is further verified through its applications in tree-based interpretable clustering and significance-based hierarchical clustering.

📄 PDF Abstract BibTeX arXiv:2507.08382

Code (0)

등록된 구현이 없습니다.

Tasks

Two-sample testing

Similar Papers 제목 키워드 기반

Testing for the appropriate level of clustering in linear regression models

2023-01-11 · James G. MacKinnon, Morten Ørregaard Nielsen, Matthew D. Webb

The overwhelming majority of empirical research that uses cluster-robust inference assumes that the clustering structure is known, even though there are often several possible ways in which a dataset could be clustered. …

Clusteringregression

Clusterability test for categorical data

2023-07-14 · Lianyu Hu, Junjie Dong, Mudi Jiang, Yan Liu 외

The objective of clusterability evaluation is to check whether a clustering structure exists within the data set. As a crucial yet often-overlooked issue in cluster analysis, it is essential to conduct such a test before…

AttributeClusteringvalid

Combining Clusters for the Approximate Randomization Test

2025-02-06 · Chun Pong Lau

This paper develops procedures to combine clusters for the approximate randomization test proposed by Canay, Romano, and Shaikh (2017). Their test can be used to conduct inference with a small number of clusters and impo…

valid

A Goodness-of-fit Test on the Number of Biclusters in a Relational Data Matrix

2021-02-23 · Chihiro Watanabe, Taiji Suzuki

Biclustering is a method for detecting homogeneous submatrices in a given observed matrix, and it is an effective tool for relational data analysis. Although there are many studies that estimate the underlying bicluster …

Clustering

Inference in IV models with clustered dependence, many instruments and weak identification

2023-06-14 · Johannes W. Ligtenberg

Data clustering reduces the effective sample size from the number of observations towards the number of clusters. For instrumental variable models I show that this reduced effective sample size makes the instruments more…