paper-with-me

홈 › Papers

Massive Data Clustering in Moderate Dimensions from the Dual Spaces of Observation and Attribute Data Clouds

2017-04-06 · Fionn Murtagh

Cluster analysis of very high dimensional data can benefit from the properties of such high dimensionality. Informally expressed, in this work, our focus is on the analogous situation when the dimensionality is moderate to small, relative to a massively sized set of observations. Mathematically expressed, these are the dual spaces of observations and attributes. The point cloud of observations is in attribute space, and the point cloud of attributes is in observation space. In this paper, we begin by summarizing various perspectives related to methodologies that are used in multivariate analytics. We draw on these to establish an efficient clustering processing pipeline, both partitioning and hierarchical clustering.

📄 PDF Abstract BibTeX arXiv:1704.01871

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeClustering

Similar Papers 제목 키워드 기반

A spectral clustering-type algorithm for the consistent estimation of the Hurst distribution in moderately high dimensions

2025-01-30 · Patrice Abry, Gustavo Didier, Oliver Orejola, Herwig Wendt

Scale invariance (fractality) is a prominent feature of the large-scale behavior of many stochastic systems. In this work, we construct an algorithm for the statistical identification of the Hurst distribution (in partic…

ClusteringModel SelectionTime Series

Differentially Private Clustering: Tight Approximation Ratios

2020-08-18 · NeurIPS 2020 12 · Badih Ghazi, Ravi Kumar, Pasin Manurangsi

We study the task of differentially private clustering. For several basic clustering problems, including Euclidean DensestBall, 1-Cluster, k-means, and k-median, we give efficient differentially private algorithms that a…

Clustering

Modified Multidimensional Scaling and High Dimensional Clustering

2018-10-24 · Xiucai Ding, Qiang Sun

Multidimensional scaling is an important dimension reduction tool in statistics and machine learning. Yet few theoretical results characterizing its statistical performance exist, not to mention any in high dimensions. B…

ClusteringDimensionality ReductionVocal Bursts Intensity Prediction

Do BERT Embeddings Encode Narrative Dimensions? A Token-Level Probing Analysis of Time, Space, Causality, and Character in Fiction

2026-04-12 · Beicheng Bei, Hannah Hyesun Chun, Chen Guo, Arwa Saghiri arxiv

Narrative understanding requires multidimensional semantic structures. This study investigates whether BERT embeddings encode dimensions of fictional narrative semantics -- time, space, causality, and character. Using an…

Massively-Parallel Heat Map Sorting and Applications To Explainable Clustering

2023-09-14 · Sepideh Aghamolaei, Mohammad Ghodsi

Given a set of points labeled with $k$ labels, we introduce the heat map sorting problem as reordering and merging the points and dimensions while preserving the clusters (labels). A cluster is preserved if it remains co…

ClusteringDimensionality Reduction