Isotropy, Clusters, and Classifiers
Whether embedding spaces use all their dimensions equally, i.e., whether they are isotropic, has been a recent subject of discussion. Evidence has been accrued both for and against enforcing isotropy in embedding spaces. In the present paper, we stress that isotropy imposes requirements on the embedding space that are not compatible with the presence of clusters -- which also negatively impacts linear classification objectives. We demonstrate this fact both mathematically and empirically and use it to shed light on previous results from the literature.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Metrics for quantifying isotropy in high dimensional unsupervised clustering tasks in a materials context
Clustering is a common task in machine learning, but clusters of unlabelled data can be hard to quantify. The application of clustering algorithms in chemistry is often dependant on material representation. Ascertaining …
ClusteringIsotropy Cliffs: The Geometric Signature of Decision-Making in Large Language Models
We investigate the geometry of decision-making in Multiple Choice Question Answering (MCQA) through the lens of isotropy. Analyzing five open-weight models across diverse datasets, we identify decision-critical transitio…
Question AnsweringIsotropy in the Contextual Embedding Space: Clusters and Manifolds
The geometric properties of contextual embedding spaces for deep language models such as BERT and ERNIE, have attracted considerable attention in recent years. Investigations on the contextual embeddings demonstrate a st…
Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations
Self-supervised speech representations can hugely benefit downstream speech technologies, yet the properties that make them useful are still poorly understood. Two candidate properties related to the geometry of the repr…
The Effect of the Intrinsic Dimension on the Generalization of Quadratic Classifiers
It has been recently observed that neural networks, unlike kernel methods, enjoy a reduced sample complexity when the distribution is isotropic (i.e., when the covariance matrix is the identity). We find that this sensit…