paper-with-me

Papers

Metrics for quantifying isotropy in high dimensional unsupervised clustering tasks in a materials context

2023-05-25 · Samantha Durdy, Michael W. Gaultois, Vladimir Gusev, Danushka Bollegala, Matthew J. Rosseinsky

Clustering is a common task in machine learning, but clusters of unlabelled data can be hard to quantify. The application of clustering algorithms in chemistry is often dependant on material representation. Ascertaining the effects of different representations, clustering algorithms, or data transformations on the resulting clusters is difficult due to the dimensionality of these data. We present a thorough analysis of measures for isotropy of a cluster, including a novel implantation based on an existing derivation. Using fractional anisotropy, a common method used in medical imaging for comparison, we then expand these measures to examine the average isotropy of a set of clusters. A use case for such measures is demonstrated by quantifying the effects of kernel approximation functions on different representations of the Inorganic Crystal Structure Database. Broader applicability of these methods is demonstrated in analysing learnt embedding of the MNIST dataset. Random clusters are explored to examine the differences between isotropy measures presented, and to see how each method scales with the dimensionality. Python implementations of these measures are provided for use by the community.

📄 PDF Abstract BibTeX arXiv:2305.16372

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing

2024-10-15 · Richard Diehl Martinez, Zebulon Goriely, Andrew Caines, Paula Buttery 외

Language models strongly rely on frequency information because they maximize the likelihood of tokens during pre-training. As a consequence, language models tend to not generalize well to tokens that are seldom seen duri…

Language ModelingLanguage ModellingSentence

Evaluating the Stability of Deep Learning Latent Feature Spaces

2024-02-17 · Ademide O. Mabadeje, Michael J. Pyrcz

High-dimensional datasets present substantial challenges in statistical modeling across various disciplines, necessitating effective dimensionality reduction methods. Deep learning approaches, notable for their capacity …

Decision MakingDeep LearningDimensionality Reduction

Interpretable Machine Learning for Spatial Science: A Lie-Algebraic Kernel for Rotationally Anisotropic Gaussian Processes

2026-05-11 · Kane Warrior, Dalia Chakrabarty arxiv

Many three-dimensional spatial fields are anisotropic, with directions of rapid and slow variation that need not align with the coordinate axes. Standard Gaussian process kernels with Automatic Relevance Determination (A…

Interpretable Machine LearningGaussian ProcessesBayesian Inference

The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models

2023-11-10 · Anton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova, Ivan Oseledets 외

In this study, we present an investigation into the anisotropy dynamics and intrinsic dimension of embeddings in transformer architectures, focusing on the dichotomy between encoders and decoders. Our findings reveal tha…

Large-Scale Evaluation of Topic Models and Dimensionality Reduction Methods for 2D Text Spatialization

2023-07-17 · Daniel Atzberger, Tim Cech, Willy Scheibel, Matthias Trapp 외

Topic models are a class of unsupervised learning algorithms for detecting the semantic structure within a text corpus. Together with a subsequent dimensionality reduction algorithm, topic models can be used for deriving…

Dimensionality ReductionSemantic SimilaritySemantic Textual SimilarityTopic Models