paper-with-me

홈 › Papers

Composite Silhouette: A Subsampling-based Aggregation Strategy

2026-04-15 · Aggelos Semoglou, Aristidis Likas, John Pavlopoulos arxiv

Determining the number of clusters is a central challenge in unsupervised learning, where ground-truth labels are unavailable. The Silhouette coefficient is a widely used internal validation metric for this task, yet its standard micro-averaged form tends to favor larger clusters under size imbalance. Macro-averaging mitigates this bias by weighting clusters equally, but may overemphasize noise from under-represented groups. We introduce Composite Silhouette, an internal criterion for cluster-count selection that aggregates evidence across repeated subsampled clusterings rather than relying on a single partition. For each subsample, micro- and macro-averaged Silhouette scores are combined through an adaptive convex weight determined by their normalized discrepancy and smoothed by a bounded nonlinearity; the final score is then obtained by averaging these subsample-level composites. We establish key properties of the criterion and derive finite-sample concentration guarantees for its subsampling estimate. Experiments on synthetic and real-world datasets show that Composite Silhouette effectively reconciles the strengths of micro- and macro-averaging, yielding more accurate recovery of the ground-truth number of clusters.

📄 PDF Abstract BibTeX arXiv:2604.13816

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Revisiting Silhouette Aggregation

2024-01-11 · John Pavlopoulos, Georgios Vardakas, Aristidis Likas

Silhouette coefficient is an established internal clustering evaluation measure that produces a score per data point, assessing the quality of its clustering assignment. To assess the quality of the clustering of the who…

Clustering

Manifold regularization based on Nystr{ö}m type subsampling

2017-10-13 · Abhishake Rastogi, Sivananthan Sampath

In this paper, we study the Nystr{\"o}m type subsampling for large scale kernel methods to reduce the computational complexities of big data. We discuss the multi-penalty regularization scheme based on Nystr{\"o}m type s…

image-classificationImage ClassificationIntrusion DetectionMulti-Task Learning+1

Globally Optimal Pose from Orthographic Silhouettes

2026-04-10 · Agniva Sengupta, Dilara Kuş, Jianning Li, Stefan Zachow arxiv

We solve the problem of determining the pose of known shapes in $\mathbb{R}^3$ from their unoccluded silhouettes. The pose is determined up to global optimality using a simple yet under-explored property of the area-of-s…

Pose Estimation

Data-Driven Subsampling in the Presence of an Adversarial Actor

2024-01-07 · Abu Shafin Mohammad Mahdee Jameel, Ahmed P. Mohamed, Jinho Yi, Aly El Gamal 외

Deep learning based automatic modulation classification (AMC) has received significant attention owing to its potential applications in both military and civilian use cases. Recently, data-driven subsampling techniques h…

Adversarial AttackAdversarial RobustnessDeep Learning

GaitASMS: Gait Recognition by Adaptive Structured Spatial Representation and Multi-Scale Temporal Aggregation

2023-07-29 · Yan Sun, Hu Long, Xueling Feng, Mark Nixon

Gait recognition is one of the most promising video-based biometric technologies. The edge of silhouettes and motion are the most informative feature and previous studies have explored them separately and achieved notabl…

Data AugmentationGait Recognition