Hierarchical Models: Intrinsic Separability in High Dimensions
It has long been noticed that high dimension data exhibits strange patterns. This has been variously interpreted as either a "blessing" or a "curse", causing uncomfortable inconsistencies in the literature. We propose that these patterns arise from an intrinsically hierarchical generative process. Modeling the process creates a web of constraints that reconcile many different theories and results. The model also implies high dimensional data posses an innate separability that can be exploited for machine learning. We demonstrate how this permits the open-set learning problem to be defined mathematically, leading to qualitative and quantitative improvements in performance.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningOpen Set LearningVocal Bursts Intensity PredictionSimilar Papers 제목 키워드 기반
Estimating the effective dimension of large biological datasets using Fisher separability analysis
Modern large-scale datasets are frequently said to be high-dimensional. However, their data point clouds frequently possess structures, significantly decreasing their intrinsic dimensionality (ID) due to the presence of …
validSeparable choices
We introduce the novel setting of joint choices, in which options are vectors with components associated to different dimensions. In this framework, menus are multidimensional, being vectors whose components are one-dime…
A Novel Intrinsic Measure of Data Separability
In machine learning, the performance of a classifier depends on both the classifier model and the separability/complexity of datasets. To quantitatively measure the separability of datasets, we create an intrinsic measur…
Generalisation and the Geometry of Class Separability
Recent results in deep learning show that considering only the capacity of machines does not adequately explain the generalisation performance we can observe. We propose that by considering the geometry of the data we ca…
Deep LearningRelative intrinsic dimensionality is intrinsic to learning
High dimensional data can have a surprising property: pairs of data points may be easily separated from each other, or even from arbitrary subsets, with high probability using just simple linear classifiers. However, thi…
Binary Classification